WRITING
Blog
Sharing insights on AI, Machine Learning, and Technology. Explore my thoughts on the latest developments and practical applications.
- The System Design Decisions Behind Big Tech Stacks — System design lessons from Uber, Netflix, Stripe, and other tech giants. A practical breakdown of distributed systems, cloud-native infrastructure, high-performance backends, and the trade-offs tech leads must understand.
- LLM Latency in Production (Part 1) — Model-Level Optimization — A tech lead's playbook for reducing LLM inference latency in production. Part 1 focuses on model-level optimization: GPU bottlenecks, memory bandwidth limits, quantization (INT8/INT4), Flash Attention, and vLLM internals.
- How to Use, Optimize and Serve an LLM in Your Production System — An end-to-end guide covering the full lifecycle of deploying LLMs in production: model selection, quantization and pruning strategies, inference optimization, and high-performance serving with vLLM and ONNX Runtime.
- LLM Latency in Production (Part 2) — Serve-Level Speed — Part 2 of the LLM latency series. Covers serve-level architecture: request batching, async queuing, load balancing, and system design patterns that stabilize P95/P99 tail latency in production LLM services.
- LLM Latency in Production (Part 3) — Engine-Level Runtime Selection — Part 3 of the LLM latency series. A deep dive into inference engine selection — vLLM, TensorRT-LLM, ONNX Runtime — and how choosing the right runtime gives you throughput, latency, and hardware efficiency for free.