Abhishek Jaiswal · Plano, TX/Open to: AI Product Manager▌
I build the AI productsI'd manage.
Five years shipping data and AI products inside a Fortune 500 financial institution — most recently a production multi-agent system running at 98% automation. On the side: an open-source framework that benchmarks 300+ LLMs, because someone has to check the models' homework.

automation reached by a production multi-agent AI system
LLMs benchmarked by my open-source eval framework
daily users on analytics products I shipped
building data & AI inside Fortune 500 financial services
Exhibit 01 · Production
Four agents, one orchestrator, zero manual QA.
Built for a Fortune 500 US financial institution (via TCS). Drop in a requirements document and the pipeline takes over — one agent extracts every policy in scope, the next writes SQL to compare the data across two Snowflake layers, a validation run grades every policy PASS / WARN / FAIL against 13 automated checks, and an analysis agent turns the results into a summary plus a QA-ready test plan you can file straight into qTest or JIRA.
I owned the product end to end: defined acceptance criteria and the quality bar, built evaluation datasets, ran LLM-as-judge scoring for accuracy and hallucination before ship, and added the governance a regulated environment demands — audit logging, access controls, human-in-the-loop on sensitive steps. It replaced a fully manual QA process and reached 98% automation with 90% fewer errors.
ORCHESTRATOR · plans · routes · retries
Requirement Doc
dropped in by the user
Policy Agent
finds every policy in scope
SQL Agent
compares two Snowflake layers
Validation Run
13 checks per policy PASS / WARN / FAIL
Analysis Agent
summary + test plan → qTest / JIRA
13 automated checks — domain rules · null checks · constraints · 1:1 mapping · derivation logic
Exhibit 02 · Open Source
Which model actually deserves your $20/month?
My LLM evaluation framework answers that with receipts: one OpenRouter key unlocks 300+ models, an LLM-as-judge scores every response across six quality dimensions, a multi-judge panel checks the judges against each other, and an agentic harness scores trajectories — not just answers. Every score traces back to a stored response, a judge rationale, and a full tool-call transcript.
$ five models · mock tool sandbox · judged on task success, tool efficiency, honesty, reasoning▌
Findings from real runs
Swapping the judge model moves composite scores by up to half a point — judge choice is a product decision, not a detail.
Pairwise model comparisons that look decisive on a leaderboard are statistical ties under bootstrap confidence intervals.
One frontier model hit a trap task honestly — searched four times, never fabricated — but also never answered. Invisible to single-turn evals.
Exhibit 03 · Open Source
An AI rules judge that proves its answers.
Rules Lawyer answers contested board-game rules questions with a verdict cited to the exact source — respecting the authority order real rulings follow. The LLM writes the explanation; it never decides which source is authoritative and never invents a ruling unsupported by retrieved text. Built with hybrid retrieval, a faithfulness judge validated against human agreement, and an abstention gate — the eval harness measures exactly how much less it hallucinates than the raw model.
“Can I play a card if I can't complete all of its effects?”
“No — every effect on the card must fully resolve, otherwise the card cannot legally be played. This is a core rule of the game.”
Confident. Fluent. Sourced from model memory — and contradicted by the game's own FAQ.
“Verdict: Yes, with limits— you resolve as much as you can. The FAQ addresses this directly and takes precedence over the rulebook's general wording.”
Every verdict cites retrieved text. When sources are silent, it abstains instead of guessing.
Illustrative exchange — the real system runs on an ingested corpus with an eval harness behind it.
Method
How I work
I don't separate “technical” from “strategic.” The best AI products get built when the person writing the Python also understands why the business needs it.
Discover
Start with the business problem, not the technology. Stakeholder interviews, metric deep-dives, and competitive research to understand what “good” looks like before a single line of code or PRD bullet.
Define
Translate ambiguous business needs into structured requirements — user stories, PRDs, and success metrics engineers can build against and stakeholders can sign off on.
Prototype
Build a working proof of concept fast — Python, LLM integrations, live UIs. Validate with real outputs, not slide decks. If it doesn't work in a notebook, it won't work in production.
Ship
Own delivery end to end in Agile sprints — from data pipeline and model integration through stakeholder UAT, validation, and go-live, for products used daily by 500+ people.
Measure
Define leading and lagging KPIs before launch, instrument tracking, and close the loop. No product is done at launch — it's done when the metrics prove it's working.
Track record
Experience
2021 — Now
Senior Analyst, Product & AI Engineering
Tata Consultancy Services · Fortune 500 financial services client · Plano, TX
- Owned roadmap and PRDs for an internal analytics platform serving Finance, Operations, and Product — cut executive time-to-insight by 30%
- Defined the evaluation and quality bar for a production multi-agent AI system shipped to 98% automation
- Led a 4-person data engineering team; drove delivery across engineering, data, and business stakeholders in Agile cycles
- Built governance and responsible-AI controls for a regulated environment — audit logging, access controls, human-in-the-loop review
2020
Data Analyst, Ad Intelligence Product
Airfind Corp · Remote, US
- Took product ownership of an early-stage ad analytics product — KPIs, A/B test methodology, behavior dashboards; recommendations drove 15% revenue growth
2017 — 2019
Research Analyst, Automation & Data Quality
Media.Net Software Services · Mumbai, India
- Built Python and SQL automation pipelines for data quality and analytics workflows — cut processing time 40%
2019 — 2020
MS, Information Technology & Analytics
Rutgers University · Newark, NJ · GPA 3.83
- Machine learning & statistics, data analysis & visualization, business analytics programming
Skills at a glance
AI Product Management · Product Strategy · PRDs · Roadmapping · User Discovery · LLM Evaluation · LLM-as-judge · Eval Pipeline Design · Agentic Systems · Multi-Agent Orchestration · RAG · Prompt Engineering · Responsible AI · Python · SQL · Snowflake · AWS · dbt · Docker · ETL & Data Pipelines · A/B Testing · KPI Design · Agile / SCRUM · Stakeholder Management
Peer-reviewed
Published research
The production work has a paper trail. Co-authored research on the same problems I ship against — data quality, lineage for auditability, and anomaly detection in insurance data. With S. Malviya and V. Koli.
Contact
Let's ship something real.
Two resumes, one builder — pick the lens that fits your role. Or grab 30 minutes on my calendar; I answer email faster than most APIs.