Available for AI projects

Kamlesh Kumar

AI Engineer · Tech Lead · Builder of LLM Products

I design and ship AI-native products — retrieval pipelines, agentic workflows and LLM-backed tooling that hold up under real traffic. Ten years of web and blockchain engineering underneath, so the models land in systems that actually stay up.

  • 5+ years shipping production software
  • LLM apps, RAG pipelines & autonomous agents
  • Blockchain, Web3 & smart contract engineering
  • Open source advocate — building in public
5+
Years building
5
Live products
10+
Devs led
Kamlesh Kumar
About

Building with AI

I started in full-stack web, spent years in blockchain and smart contracts, and now spend most of my time on AI systems. That path matters: the hard part of an LLM product is rarely the prompt — it's retrieval quality, evaluation, cost, latency and the failure modes nobody writes about in a demo. I build the whole thing, from the retrieval layer to the UI, and I ship it publicly.

LLM Application Engineering

Production apps on top of Claude, GPT and open-weight models — prompt architecture, structured output, evals, token and latency budgets, graceful fallback when a provider goes sideways.

RAG & Retrieval Systems

Chunking strategies, embeddings, vector stores and hybrid search that actually returns the right passage. Grounded answers with citations instead of confident nonsense.

Agents & Automation

Tool-calling agents that do real work: multi-step workflows, MCP servers, human-in-the-loop checkpoints, and the guardrails that keep an autonomous loop from running away.

AI Infrastructure & Safety

Streaming APIs, caching, cost observability, secret hygiene and prompt-injection defence — the unglamorous layer that decides whether an AI feature survives contact with users.

Research & Notes

What I'm working through

Open questions I'm actively chewing on, and the writing that comes out of it.

Open Questions

  1. 01
    How do you evaluate a RAG system without a labelled dataset?
    Active

    Most retrieval work starts with no ground truth. Interested in bootstrapping evals from production traffic, LLM-as-judge with calibration checks, and knowing when a retrieval score is measuring anything real.

  2. 02
    What guardrails actually stop an agent loop from running away?
    Active

    Step budgets and tool allowlists are the easy part. The harder question is detecting when an agent is confidently wrong mid-run, and where a human checkpoint belongs without destroying the workflow.

  3. 03
    Where does prompt injection defence belong in the stack?
    Exploring

    Input filtering, output constraints, tool-permission boundaries — each catches a different class. Working through which layer earns its complexity for an app that reads untrusted content.

  4. 04
    What is the real cost curve of an LLM feature at scale?
    Background

    Token spend is visible; latency budgets, cache hit rates and the retries hiding behind a p99 are not. Building the observability to answer this before the invoice does.

Learning

Currently studying

The field moves weekly. Here's where my attention actually is right now.

Evaluation & benchmarking for LLM systems

80%

Building repeatable evals: golden sets, LLM-as-judge with human spot-checks, and regression suites that catch a prompt change before users do.

Agent architectures & tool protocols

70%

Multi-step planning, MCP servers, tool-permission design, and the failure modes that only appear once an agent runs unattended.

Vector search & hybrid retrieval

75%

Embedding choice, chunking strategy, reranking, and combining dense with keyword search when neither wins alone.

Applied model fine-tuning

45%

LoRA and adapter approaches — mostly to develop judgement about when fine-tuning beats better retrieval, which is less often than it looks.

Toolkit

My Tech Arsenal

What I reach for, grouped by what it actually does.

AI & Machine Learning

Claude API
OpenAI API
LangChain
RAG / Vector DBs
Hugging Face
PyTorch
TensorFlow
scikit-learn
Pandas
NumPy

Web & Backend

Next.js
React
TypeScript
JavaScript
Node.js
Python
FastAPI
Go
GraphQL

Blockchain & Web3

Ethereum
Solidity

Data & Infrastructure

PostgreSQL
MongoDB
Redis
SQL/NoSQL
AWS
GCP
Vercel
Docker
Kubernetes
Git
Linux
Writing

Latest from X

Threads and articles on AI engineering, LLM patterns and building in public.

Journey

Experience

Founder & AI Engineer2024 - Present
Independent Products
Designing, building and running AI-native products end to end — ForgeLearn, SecureEnv and RoastMyProd. LLM orchestration, retrieval pipelines, evals and cost control, plus everything around them: auth, billing, infra and support.
LLM AppsRAGAgentsNext.jsTypeScriptPythonProduct
Tech Lead & Architect2021 - Present
Advantev Solutions
Spearheading architecture and development for AI, Blockchain, and SaaS projects. Leading a cross-functional team to deliver enterprise-grade systems and innovative digital solutions.
AI/MLBlockchainPythonReactTypeScriptAWSTeam Leadership
Contributor / Mentor2017 - Present
Open Source & Community
Actively contributing to open-source projects (AI, Web3, Next.js, Solidity). Mentoring developers, giving talks, and building tools for the community.
Open SourceNext.jsSolidityTypeScriptMentorshipTalks
Say hello

Contact & Social

Kamlesh Kumar
Kamlesh Kumar
AI Engineer · Tech Lead
Let's connect! Whether you're building something with LLMs, want to collaborate, or just want to argue about agent design — reach out on any platform below.