Overview
Social Links
About
Hello, World! I'm Ketan Kumar, an Associate Software Engineer at AIonOS, working on backend systems in Python and Rust. I focus on async concurrency and LLM infrastructure, and I build real-time systems on Tokio.
At AIonOS, I built the concurrency layer of a FastAPI microservice that converts airline fare rules into validated, structured XML, sustaining 1,500+ documents/minute through bounded semaphores and a thread pool that keeps blocking work off the event loop. I've also hardened XML-parsing endpoints against XXE attacks and built a provider-agnostic LLM client for OpenAI, vLLM, and Ollama.
Outside of work, I build real-time systems in Rust — from a crypto arbitrage scanner aggregating prices across 14 liquidity sources to a Solana DePIN platform with on-chain stake economics.
Feel free to reach out if you're interested in collaborating!
Experience
LLM platform that converts airline fare rules (ATPCO categories 16, 31, 33) into validated, structured XML.
- Built the concurrency layer of a FastAPI microservice sustaining 1,500+ fare-rule documents/minute: separate asyncio semaphores cap in-flight requests (100) and LLM calls (50) to decouple request intake from model rate limits, and a bounded ThreadPoolExecutor keeps blocking XML work off the event loop.
- Added a request-scoped LLM response cache keyed on a content hash of the parsed rule text, so legs and fare-basis codes producing identical prompts in a multi-leg request share one LLM call.
- Wrote automated cross-category mismatch detection and correction for CAT16/31/33 output, fixing or retrying malformed LLM responses before results are persisted asynchronously to MongoDB.
- Hardened every XML-parsing endpoint against XXE and entity-expansion attacks by migrating to defusedxml, resolving all CodeQL findings (0 open alerts) and locking the fix in with 10 regression tests.
- Built a provider-agnostic LLM client for OpenAI, self-hosted vLLM, and Ollama with retry and backoff, and instrumented per-call token usage, latency, and queue time in New Relic for observability.