SDE Intern (ML + Dev)
Rimo LLC, Japan • Remote
- Architected and shipped RAG chat end-to-end — a production meeting transcript search system serving 800K users — on a 3-service microarchitecture (Go/Python + Cloud Run Jobs), with full ownership from design to deployment monitoring.
- Built a real-time SSE streaming pipeline combining BM25 + Dense vector search (Ruri embeddings), RRF, and RuriV3 reranking; reduced query latency by 12% handling ~6,500 req/sec in production.
- Designed an async distributed indexing system on GCP Cloud Run Jobs (NVIDIA L4 GPU) for chunking, embedding generation, and MongoDB Atlas writes — decoupled from the API layer to eliminate resource contention at scale.
