Harshit Malik
AI engineer, LLM inference and multi-agent systems
I like making models fast and cheap to serve, and building agents that hold up under real traffic.
Current Focus
I’m deeply interested in inference economics, the arithmetic of what it actually costs to serve a model: where every gigabyte of GPU memory goes, what a millisecond of latency is worth, and why the flag everyone tunes is rarely the one that moves the number. Memory, latency and cost are remainders, not settings; you don’t configure a serving system, you derive it.
I believe the next step for agents is systems that are engineered, not just prompted: multi-agent orchestration with deterministic validation, tool use that survives long horizons, and voice pipelines where the latency budget is a human conversation, not a benchmark.
“My motto is, if you want to win the lottery, you have to make the money to buy a ticket.”
Experience
TheAgentic – AI Engineer, 2026–. CAPAtrace: multi-agent orchestration with deterministic validation and verifiable audit trails for regulated document review, cutting review time by 97%.
StatusNeo – Generative AI Engineer, 2025–26. An LLM platform that compiled plain English into Playwright suites using Claude, MCP and DOM diffing, and a LangChain workflow engine running 50K+ automated tests a day in CI.
Misty Interactive – Software Engineer Intern, 2025. A retrieval-backed assistant wired into CRM, on a backend scaled to 15K concurrent users at sub-150 ms p95.
Recent writing
- Inside the KV Cache: The Life of a Gigabyte 29 Sep 2026
Projects
- Blokrly – a strict website blocker for Chrome, with scheduled focus sessions.
- OpenRepo – RAG over arbitrary repos.
- vLLM KV cache calculator – will your model boot, and how many tokens of cache will you get.