Mehedi Hasan — AI/ML Engineer
Senior AI/ML Engineer with 5+ years of experience building production AI systems across LLMs, RAG, document intelligence, and distributed services.
Selected Work
- Agami Assistant - Voice-First Enterprise AI Assistant — Voice-first enterprise assistant providing grounded spoken answers from product catalogs, policies, and IT records, with live transcripts and captions.
- Search Microservice - High-Performance Product Discovery — Typo-tolerant product discovery for 1M products and ~29K brands; reached 179 QPS at 3.17 ms P95 for warm traffic and 1,046 QPS under concurrent load.
- Shorol Notes — AI-Powered Note-Taking — Voice notes with AI transcription and summarization; calendar sync with Bangla-first UX. Pluggable backends supporting OpenAI, Ollama, and Hugging Face models.
- Japanese Lawyer Assistant — Legal Q&A over a Japanese corpus using RAG, FAISS retrieval, and an instruct LLM with a knowledge-base pipeline and custom prompts.
- Outlet Fraud Detection - No-Label Verification Screening — No-label fraud screening for outlet-verification photos. Every image is compared only against other photos from the same outlet — visually isolated shots and reused uploads get flagged with an evidence-based reason.
- Aethion RAG - Naive vs Streaming Retrieval Shootout — Full-stack AI agent answering questions over company documents, comparing Naive RAG against StreamRAG — retrieval fired in parallel with generation and streamed to the user.
- AI Research Agent - Multi-Source LLM Assistant — LLM research assistant and multi-source AI agent capable of answering complex research questions across sources.
- ECG Arrhythmia Classification - 2-D CNN — 2-D convolutional classifier detecting arrhythmia from ECG beats through careful signal preprocessing and feature extraction.
- ToLet Dhaka - Rental Listings — Flat-renting website with property posts and contact details, enhanced with interactive elements.
- School Management System — School administration console for student records, classes, and fees.
- Bangla Alarm Clock — Alarm clock application with a Bangla interface.
- Dhaka City Bus Routes — Route-finder Android app covering Dhaka city transit lines with map lookup.
- Bank Management System — Transactional banking features with account handling and balance tracking.
- DX Ball Remastered — Classic brick-breaker ball and paddle game with 4 new levels.
- Hospital Management — Full-stack hospital management system with patient, doctor, and appointment workflows.
- Railway Station Management System — Railway station records and scheduling with client-server components in PL/SQL.
- BanglaNet - Handwritten Recognition — Lightweight CNN recognizing 50 Bangla handwritten characters and 10 numerals for OCR-style digitization.
Writing
- The System Design Decisions Behind Big Tech Stacks — System design lessons from Uber, Netflix, Stripe, and other tech giants. A practical breakdown of distributed systems, cloud-native infrastructure, high-performance backends, and the trade-offs tech leads must understand.
- LLM Latency in Production (Part 1) — Model-Level Optimization — A tech lead's playbook for reducing LLM inference latency in production. Part 1 focuses on model-level optimization: GPU bottlenecks, memory bandwidth limits, quantization (INT8/INT4), Flash Attention, and vLLM internals.
- How to Use, Optimize and Serve an LLM in Your Production System — An end-to-end guide covering the full lifecycle of deploying LLMs in production: model selection, quantization and pruning strategies, inference optimization, and high-performance serving with vLLM and ONNX Runtime.
- LLM Latency in Production (Part 2) — Serve-Level Speed — Part 2 of the LLM latency series. Covers serve-level architecture: request batching, async queuing, load balancing, and system design patterns that stabilize P95/P99 tail latency in production LLM services.
- LLM Latency in Production (Part 3) — Engine-Level Runtime Selection — Part 3 of the LLM latency series. A deep dive into inference engine selection — vLLM, TensorRT-LLM, ONNX Runtime — and how choosing the right runtime gives you throughput, latency, and hardware efficiency for free.