Ketan Kumar's avatar
Ketan Kumar 
hello
GitHub
Ketan Kumar's avatar
text-3xl text-zinc-950 font-semibold

Ketan Kumar 

Creating with code, driven by passion.

Overview

Social Links

About

Hello, World! I'm Ketan Kumar, an Associate Software Engineer at AIonOS, working on backend systems in Python and Rust. I focus on async concurrency and LLM infrastructure, and I build real-time systems on Tokio.

At AIonOS, I built the concurrency layer of a FastAPI microservice that converts airline fare rules into validated, structured XML, sustaining 1,500+ documents/minute through bounded semaphores and a thread pool that keeps blocking work off the event loop. I've also hardened XML-parsing endpoints against XXE attacks and built a provider-agnostic LLM client for OpenAI, vLLM, and Ollama.

Outside of work, I build real-time systems in Rust — from a crypto arbitrage scanner aggregating prices across 14 liquidity sources to a Solana DePIN platform with on-chain stake economics.

Feel free to reach out if you're interested in collaborating!

Experience

LLM platform that converts airline fare rules (ATPCO categories 16, 31, 33) into validated, structured XML.

  • Built the concurrency layer of a FastAPI microservice sustaining 1,500+ fare-rule documents/minute: separate asyncio semaphores cap in-flight requests (100) and LLM calls (50) to decouple request intake from model rate limits, and a bounded ThreadPoolExecutor keeps blocking XML work off the event loop.
  • Added a request-scoped LLM response cache keyed on a content hash of the parsed rule text, so legs and fare-basis codes producing identical prompts in a multi-leg request share one LLM call.
  • Wrote automated cross-category mismatch detection and correction for CAT16/31/33 output, fixing or retrying malformed LLM responses before results are persisted asynchronously to MongoDB.
  • Hardened every XML-parsing endpoint against XXE and entity-expansion attacks by migrating to defusedxml, resolving all CodeQL findings (0 open alerts) and locking the fix in with 10 regression tests.
  • Built a provider-agnostic LLM client for OpenAI, self-hosted vLLM, and Ollama with retry and backoff, and instrumented per-call token usage, latency, and queue time in New Relic for observability.
PythonFastAPIasyncioMongoDBLLM IntegrationdefusedxmlCodeQLNew Relic

Projects

Tech Stack

Python iconRust iconTypeScript iconJavaScript iconFastAPI iconNode.js iconExpress.js iconSolana iconAnchor iconOpenAI API iconReact iconPostgreSQL iconMySQL iconMongoDB iconRedis iconDocker iconKubernetes iconPrisma iconGit iconLinux icon