AI that shows its work

Rohit Agrawal — principal-level engineer, fourteen years across distributed systems and production reliability, now building independent AI systems that cite their sources, disclose their uncertainty, refuse what they cannot support, and gate their own releases.

Experience
14+ yrs — Oracle, Amazon, LimeRoad, Mobileum, Snapdeal, Subex
Most recent
Oracle Principal MTS · 2019–2026
Since April 2026
Building independent AI systems, full time
Seeking
Senior / Principal — AI Platform · LLMOps · Forward-Deployed AI · AI Quality Engineering Architect · AI Test Automation & Agentic AI Leader
Base
Bengaluru, India · IST (UTC+5:30)
Relocation
Open worldwide

CiteVyn, golden case claude_api_006

Refused

“What is the capital of France?”

citations
0
domain
unsupported
confidence
none
answered
No — by policy
It plainly knows the answer and refuses anyway, because no indexed source supports it. Verbatim from citevyn/backend/artifacts/golden_report.json, recorded 17 July 2026.
Live — cold-starts

CiteVyn

“Can I trust this answer — and trace every claim to its source?”

Citation-grounded Q&A over official AI documentation. Answers quote their sources verbatim; where no source supports an answer, CiteVyn refuses instead of guessing. Index updates reach production only through an evaluation gate.

Golden regression run

52 / 52 passed

A red case blocks the index from reaching production.

cases
52
failed
0
run
17 Jul 2026
gate
Blocks promotion
citevyn/backend/artifacts/golden_report.json at df8cfc3. Read from the recorded run, not restated.
Answers Quoted verbatim
Citations Every claim
Refusal A feature
Release gate 52 golden cases
Tests 1,036 defined
Stack FastAPI · pgvector
Live

Quorum‑AI

“What happens when four models disagree about your question?”

One question runs against four models in parallel; they critique one another for two rounds, and a synthesis returns consensus, disagreement, source support, uncertainty, and a recommendation. The cost is approved before anything runs, and any fallback or simulation is disclosed, never hidden.

Production readiness review

Go

The review before it said No-Go. That one is kept too.

decision
Go — single instance
dated
21 Jun 2026
scope
MVP, small user base
superseded
No-Go of 16 Jun
quorum-ai/docs/95-production-readiness-review.md at 8ca6a98. Both decisions are in the file; neither was removed.
Models 4 in parallel
Debate 2 critique rounds
Verdict 5 fields
Cost Approved first
Memory None — ephemeral
Deployed — not answering

SaafSaans

“Is it safe for me to go outside right now — and if not, when?”

A Delhi-NCR air-quality companion that scores your risk — age, condition, planned activity — rather than the city's average, across 21 stations, and answers questions with cited health guidance. Every mode is labelled: live, deterministic fallback, or sample. The Hindi draft ships behind a banner saying no Hindi speaker has reviewed it yet.

Entered at Build with AI (Elastic × GDG Cloud New Delhi, 18 July 2026) as a four-tab Streamlit app that already existed, then rebuilt over the next three days into what runs today. It lives on one small machine that scales to zero when idle. Right now that machine is not answering — the address resolves and the server accepts the connection, then closes it without sending anything. The fault is being investigated. This page says so rather than leaving you to find out by clicking.

Measured at head, today

Counted, not claimed

Every number here has a command that checks it.

test functions
628
test files
32
seeded advisories
43
commits
161
saaf-saans at 5e5037a, counted on 8 Aug 2026. Its own case study lists 25 files and 117 commits — true at an older commit, and stale now. That is why these were re-counted rather than copied.
Risk Per person
Scale CPCB bands
Timing Best window
Coverage 21 stations
Hindi Gated — unreviewed
Phase 1 — No-Go

NarraTwin AI

“Can project knowledge become a walkthrough without inventing a claim?”

Grounded walkthrough generation with citations, claim evaluation, consent checks, and release gates that run before anything is generated. Its own release-readiness review currently reads No-Go — so it is not deployed, and this page says so.

It is shown anyway, because the gate holding is the point.

Release readiness review

No-Go

No-Go for production release. No release tag has been created.

dated
1 Jul 2026
release tags
0
blocked
Paid providers, video export
allowed
Local mock demo
narratwin docs/RELEASE_READINESS_REVIEW.md at 2ce5731. The repo has two tags; both mark stages, neither is a release.
Gates Claim · consent · release
Deployment None — local only
Distribution Withheld by design
In progress — closed

Carried, not shown

Two systems are still being built. They stay closed until they can be judged on finished work rather than on intent. They are named because a record that discloses its gaps should also say what exists — and the question each one is built to answer costs nothing to state.

EvalAxis

“Why can a failing test stop a release, when a measured drop in answer quality cannot?”

Evaluates LLM, RAG, and agent changes with evidence — and blocks CI on a quality regression.

In progress · closed

Aegis Contracts

“What should one AI system be allowed to promise another — and who checks?”

Early-stage work on contract-shaped guarantees between AI systems.

In progress · closed
Two zipped garment bags on a rail, each carrying a name tag. EVALAXIS AEGIS-CONTRACTS

One message away

Open to senior and principal roles in Forward-Deployed AI, AI Platform Engineering, and LLMOps / AI Reliability — and closely aligned AI quality and platform work.

Based
Bengaluru, India · IST (UTC+5:30)
Open to
Global relocation · international travel