Bhuvanesh Selvaraj
I write about AI infrastructure, agents, and inference systems. How models actually serve requests, how agent loops behave once they leave the demo, and what any of it means for the people building on top.
Notes
all →- Anatomy of an inference requestWhat happens between hitting enter and seeing the first token, and why the two numbers that describe it are set by completely different machinery.
- The agent loop in productionThe loop itself is trivial. Everything that makes an agent good or bad in production lives in the parts the diagram hides.
- Context engineering is memory managementThe context window is not a text box. It is scarce, contended memory, and the old disciplines of managing scarce memory apply almost unchanged.
Ideas
all →Also here: small tools I have shipped.