Kamlesh Kumar
AI Engineer · Tech Lead · Builder of LLM Products
I design and ship AI-native products — retrieval pipelines, agentic workflows and LLM-backed tooling that hold up under real traffic. Ten years of web and blockchain engineering underneath, so the models land in systems that actually stay up.
- 5+ years shipping production software
- LLM apps, RAG pipelines & autonomous agents
- Blockchain, Web3 & smart contract engineering
- Open source advocate — building in public

Building with AI
I started in full-stack web, spent years in blockchain and smart contracts, and now spend most of my time on AI systems. That path matters: the hard part of an LLM product is rarely the prompt — it's retrieval quality, evaluation, cost, latency and the failure modes nobody writes about in a demo. I build the whole thing, from the retrieval layer to the UI, and I ship it publicly.
LLM Application Engineering
Production apps on top of Claude, GPT and open-weight models — prompt architecture, structured output, evals, token and latency budgets, graceful fallback when a provider goes sideways.
RAG & Retrieval Systems
Chunking strategies, embeddings, vector stores and hybrid search that actually returns the right passage. Grounded answers with citations instead of confident nonsense.
Agents & Automation
Tool-calling agents that do real work: multi-step workflows, MCP servers, human-in-the-loop checkpoints, and the guardrails that keep an autonomous loop from running away.
AI Infrastructure & Safety
Streaming APIs, caching, cost observability, secret hygiene and prompt-injection defence — the unglamorous layer that decides whether an AI feature survives contact with users.
What I'm working through
Open questions I'm actively chewing on, and the writing that comes out of it.
Open Questions
- 01
How do you evaluate a RAG system without a labelled dataset?
ActiveMost retrieval work starts with no ground truth. Interested in bootstrapping evals from production traffic, LLM-as-judge with calibration checks, and knowing when a retrieval score is measuring anything real.
- 02
What guardrails actually stop an agent loop from running away?
ActiveStep budgets and tool allowlists are the easy part. The harder question is detecting when an agent is confidently wrong mid-run, and where a human checkpoint belongs without destroying the workflow.
- 03
Where does prompt injection defence belong in the stack?
ExploringInput filtering, output constraints, tool-permission boundaries — each catches a different class. Working through which layer earns its complexity for an app that reads untrusted content.
- 04
What is the real cost curve of an LLM feature at scale?
BackgroundToken spend is visible; latency budgets, cache hit rates and the retries hiding behind a p99 are not. Building the observability to answer this before the invoice does.
Writing
- Collectionkaundal.vip
Notes on AI engineering, LLM patterns and building in public
The main writing archive: working notes on retrieval quality, evaluation, prompt architecture and what breaks when an LLM feature meets real users.
- Collectionsolidity.today
Solidity, smart contracts and Web3 engineering
Long-form pieces on contract design, auditing habits and the practical side of shipping on-chain systems.
- Collection@kkworld
Threads and articles on X
Shorter, faster notes — things learned mid-build, benchmarks worth sharing, and arguments about agent design.
Currently studying
The field moves weekly. Here's where my attention actually is right now.
Evaluation & benchmarking for LLM systems
80%Building repeatable evals: golden sets, LLM-as-judge with human spot-checks, and regression suites that catch a prompt change before users do.
Agent architectures & tool protocols
70%Multi-step planning, MCP servers, tool-permission design, and the failure modes that only appear once an agent runs unattended.
Vector search & hybrid retrieval
75%Embedding choice, chunking strategy, reranking, and combining dense with keyword search when neither wins alone.
Applied model fine-tuning
45%LoRA and adapter approaches — mostly to develop judgement about when fine-tuning beats better retrieval, which is less often than it looks.
My Tech Arsenal
What I reach for, grouped by what it actually does.
AI & Machine Learning
Web & Backend
Blockchain & Web3
Data & Infrastructure
Things I've Built
Products I design, build and run myself — all live in production.
ForgeLearn
A hands-on learning platform for developers — practical, project-shaped lessons instead of passive video, with AI assistance guiding you through each build.
SecureEnv
Secrets and environment-variable security for teams — scan, store and share config safely so credentials stop leaking through .env files, chat threads and commits.
RoastMyProd
Drop in your product or landing page and get an unfiltered AI critique — positioning, copy, UX and conversion gaps, delivered as blunt, actionable feedback.
Pro Kaundal
A growing collection of free developer tools — no signup, no paywall, no credit card. Just open the one you need and use it.
Kaundal VIP — Blog
Where I write it all down: AI engineering, LLM and RAG patterns, Web3, and hard-won notes from building and shipping side projects in public.
Latest from X
Threads and articles on AI engineering, LLM patterns and building in public.
Experience
Contact & Social
