Back

Bikash Lama

Senior Software Engineer & Project Lead

Technical Project Manager with 8+ years of experience delivering web services and AI-powered SaaS products in Japan. Experienced in leading cross-functional teams, project planning, stakeholder coordination, technical decision-making, and end-to-end delivery using Ruby on Rails, Python, Vue.js, AWS, Terraform, and Claude API.

8+ yrs full-stack delivery 20-person team lead Production RAG · verified citations LLM harness engineering 10k+ monthly transactions ~100% uptime SLA · 2 years
Featured Projects
Restaurant Management SaaSClient SaaS · 2 years · ongoing
Lead Engineer · Cloco

Grew a small MVP into a multi-tenant SaaS over two years of continuous delivery — now integrated into an ecosystem where the same restaurant is discoverable and reservable across collaborating partner platforms.

  • Google reservation flow + cross-platform listing — one restaurant is bookable from any partner in the ecosystem.
  • AI-in-the-loop optimization: used Claude and Cursor to surface API bottlenecks and tighten hot-path UI renders.
  • Multi-OEM architecture deployed via Terraform with a dev → staging → production CI/CD gate.
Stack: AI Tooling · Ruby on Rails · Vue.js · AWS · Terraform
AI document translator (JP → EN)Internal tool · 2026
Lead Engineer · Cloco

An internal tool that translates any Japanese document — Excel, Word, drawio diagrams — to English via Claude on the fly.

  • Format-preserving extraction across xlsx, docx, drawio — formulas, styles, and shape positions stay intact; only the human-readable text is replaced.
  • Translation cache keyed on normalized source text + target language: identical strings are never paid for twice.
  • Tiered model routing — bulk and low-complexity content goes through cheaper models; technical or mixed-domain text routes to the larger one.
Stack: AI Tooling · Claude API · Caching
Live commerce & consultation platformIn-house product · 2020—2022
Full-stack Engineer · SPINSHELL

A WebRTC-based real-time consultation product that lets customers talk live to agents across sectors — banking advisory, telemedicine, retail product demos. COVID pushed demand sharply, forcing a hard-deadline reengineering of the thin MVP into a production-grade architecture without slowing the delivery cadence.

  • Real-time video consultation over WebRTC + Twilio, shipped across banking advisors, online clinics, and apparel retail under one product surface.
  • Re-engineered the MVP skeleton from the ground up as customer load grew — moved from 'demo works' to '100% uptime SLA' shape.
  • Observability pipeline tuned for live-call triage: structured logs make it possible to chase a dropped call back to the exact request that failed.
Stack: Python · Twilio · WebRTC · Vue.js · AWS
Personal Projects — AI Experiments
Internal AI assistant — RAG with citation verificationPersonal project · 2026
Solo build · learning

A chatbot for internal document Q&A. The LLM writes the answer; my code checks every citation before the user sees it. Ungrounded answers get replaced with a refusal server-side.

  • Streamed answers get replaced with a refusal if a parser can't tie every [N] citation to a real retrieved chunk.
  • Refusal short-circuits before the LLM even runs when retrieval returns nothing above threshold — no tokens spent, no hallucinated citations.
  • Per-user daily token caps enforced in application code. Every OpenAI call writes cost to a `usage_events` table for the admin dashboard.
Stack: Next.js 15 · OpenAI API · pgvector · Postgres · TypeScript
LLM harness pattern — eight rules for keeping AI honestPersonal project · 2026
Solo build · learning

A Python script that drives a real browser via plain-English instructions. First draft was 80 lines and lied about everything. Rebuilding it produced a reusable pattern I now apply to every piece of LLM-touching code.

  • Every tool call is validated by code; success is proven by re-reading state, not by the LLM's claim.
  • Cost and iteration ceilings live in the harness — the LLM can't loop indefinitely on a broken action.
  • Refusals short-circuit in code, not by prompt persuasion. Same pattern powers the RAG assistant's citation check.
Stack: Python · Claude API · Browser automation · LLM harness
Experience
Senior Software Engineer · Project LeadNov 2022 — Present
Cloco Inc. Tokyo, Japan · Hybrid
  • Full-stack engineering on the Restaurant Management SaaS — Ruby on Rails backend, Vue.js + Vuex frontend, AWS infrastructure — owning features end-to-end from data model and API design through UI implementation, testing, and deploy.
  • Coordinated a 20-person team (engineers + QA) on the multi-tenant client SaaS — owning sprint cadence, release planning, feature scope, and the cross-partner integrations (Google reservation flow, cross-platform listing) that make the ecosystem work.
  • Drove team-wide AI-augmented development — integrated Claude and Cursor into the daily workflow and codified a shared code skeleton and conventions, so velocity grows without compromising review rigor or post-incident traceability.
  • Built an internal AI document translator (Claude API) that turns Japanese xlsx / docx / drawio files into English while preserving formulas, styles, and shape layouts — cache + tiered model routing keep cost bounded under heavy use.
  • Standardized multi-tenant SaaS deploys with Terraform — dev, staging, and production go through the same modules; cutovers are debuggable and reversible.
Full Stack EngineerSep 2020 — Oct 2022
SPINSHELL INC. Tokyo, Japan
  • Led the rebuild of an in-house live commerce & consultation platform (WebRTC + Twilio) from MVP to production-grade as COVID drove demand sharply across banking, telemedicine, and apparel retail.
  • Designed a version-retention deploy strategy that lets releases ship without dropping in-flight video calls — sustained near-100% availability across multi-sector clients over a 2-year run.
  • Built observability tuned for live-call triage; integrated Stripe/PayPal payment flows and Twilio/WebRTC video infrastructure, powering 10k+ monthly transactions across the product.
  • Designed PubSub/RPC messaging and JWT-secured public REST APIs for partner integrations.
Dispatch Web EngineerSep 2019 — Sep 2020
PERSOL Technology Staff Co., Ltd. Tokyo, Japan
  • Dispatched to Rakuten's e-commerce team on a short engagement — built Python pipelines that mined trending Twitter terms and surfaced them as market signals.
  • Output fed product and SEO decisions: which categories were heating up, which keywords to target, which themes the site should lean into.
Junior Software Developer2017 — 2019
GBD · E-multitech Nepal
  • Two years on an e-commerce product as a junior engineer — shipped bug fixes and feature work under the team lead's direction, learning the day-to-day rhythm of keeping a live product moving.
Skills
Languages: Python, Go, JavaScript, Ruby, SQL
Frameworks: Ruby on Rails, Vue.js, Vuex, AngularJS
Cloud & Infra: AWS, Terraform, Lambda, CloudWatch
Realtime & APIs: WebSocket, PubSub / RPC, REST, JWT
Practices: CI/CD, Git, TeamCity, Agile
AI: Claude, OpenAI API, Cursor, RAG, LLM harness, Prompt Eng., pgvector
Education
Bachelor In Information ManagementTribhuvan University2013 — 2017
High SchoolMoonlight Higher Secondary School2011-2013