Abhishek Jaiswal · Plano, TX/Open to: AI Product Manager

I build the AI productsI'd manage.

Five years shipping data and AI products inside a Fortune 500 financial institution — most recently a production multi-agent system running at 98% automation. On the side: an open-source framework that benchmarks 300+ LLMs, because someone has to check the models' homework.

Scroll
0%

automation reached by a production multi-agent AI system

0+

LLMs benchmarked by my open-source eval framework

0+

daily users on analytics products I shipped

0yrs

building data & AI inside Fortune 500 financial services

Exhibit 01 · Production

Four agents, one orchestrator, zero manual QA.

Built for a Fortune 500 US financial institution (via TCS). Drop in a requirements document and the pipeline takes over — one agent extracts every policy in scope, the next writes SQL to compare the data across two Snowflake layers, a validation run grades every policy PASS / WARN / FAIL against 13 automated checks, and an analysis agent turns the results into a summary plus a QA-ready test plan you can file straight into qTest or JIRA.

I owned the product end to end: defined acceptance criteria and the quality bar, built evaluation datasets, ran LLM-as-judge scoring for accuracy and hallucination before ship, and added the governance a regulated environment demands — audit logging, access controls, human-in-the-loop on sensitive steps. It replaced a fully manual QA process and reached 98% automation with 90% fewer errors.

ORCHESTRATOR · plans · routes · retries

Requirement Doc

dropped in by the user

Policy Agent

finds every policy in scope

SQL Agent

compares two Snowflake layers

Validation Run

13 checks per policy PASS / WARN / FAIL

Analysis Agent

summary + test plan → qTest / JIRA

13 automated checks — domain rules · null checks · constraints · 1:1 mapping · derivation logic

PASSWARNFAIL
Claude SonnetSnowflake MCPPythonAgent OrchestrationLLM-as-judge EvaluationqTest · JIRA

Exhibit 02 · Open Source

Which model actually deserves your $20/month?

My LLM evaluation framework answers that with receipts: one OpenRouter key unlocks 300+ models, an LLM-as-judge scores every response across six quality dimensions, a multi-judge panel checks the judges against each other, and an agentic harness scores trajectories — not just answers. Every score traces back to a stored response, a judge rationale, and a full tool-call transcript.

github.com/abhishek2395/llm-eval
agentic-harness · trajectory scores · live run jul 2026
claude-sonnet-4.5
5.00
deepseek-v3
5.00
gpt-4o-mini
5.00
gemini-3.5-flash
4.55
llama-3.3-70b
3.65

$ five models · mock tool sandbox · judged on task success, tool efficiency, honesty, reasoning

Findings from real runs

±0.5

Swapping the judge model moves composite scores by up to half a point — judge choice is a product decision, not a detail.

12/15

Pairwise model comparisons that look decisive on a leaderboard are statistical ties under bootstrap confidence intervals.

0 → ∅

One frontier model hit a trap task honestly — searched four times, never fabricated — but also never answered. Invisible to single-turn evals.

Exhibit 03 · Open Source

An AI rules judge that proves its answers.

Rules Lawyer answers contested board-game rules questions with a verdict cited to the exact source — respecting the authority order real rulings follow. The LLM writes the explanation; it never decides which source is authoritative and never invents a ruling unsupported by retrieved text. Built with hybrid retrieval, a faithfulness judge validated against human agreement, and an abstention gate — the eval harness measures exactly how much less it hallucinates than the raw model.

FAQoverridesErrataoverridesRulebook
github.com/abhishek2395/Rules-Lawyer

“Can I play a card if I can't complete all of its effects?”

Raw model no citation

“No — every effect on the card must fully resolve, otherwise the card cannot legally be played. This is a core rule of the game.”

Confident. Fluent. Sourced from model memory — and contradicted by the game's own FAQ.

Rules Lawyer grounded

“Verdict: Yes, with limits— you resolve as much as you can. The FAQ addresses this directly and takes precedence over the rulebook's general wording.”

cite: FAQ p.4context: Rulebook §5

Every verdict cites retrieved text. When sources are silent, it abstains instead of guessing.

Illustrative exchange — the real system runs on an ingested corpus with an eval harness behind it.

RAGHybrid Retrieval + RerankDeterministic PrecedenceFaithfulness JudgeAbstention GateQdrantClaude Opus

Method

How I work

I don't separate “technical” from “strategic.” The best AI products get built when the person writing the Python also understands why the business needs it.

01

Discover

Start with the business problem, not the technology. Stakeholder interviews, metric deep-dives, and competitive research to understand what “good” looks like before a single line of code or PRD bullet.

02

Define

Translate ambiguous business needs into structured requirements — user stories, PRDs, and success metrics engineers can build against and stakeholders can sign off on.

03

Prototype

Build a working proof of concept fast — Python, LLM integrations, live UIs. Validate with real outputs, not slide decks. If it doesn't work in a notebook, it won't work in production.

04

Ship

Own delivery end to end in Agile sprints — from data pipeline and model integration through stakeholder UAT, validation, and go-live, for products used daily by 500+ people.

05

Measure

Define leading and lagging KPIs before launch, instrument tracking, and close the loop. No product is done at launch — it's done when the metrics prove it's working.

Track record

Experience

2021 — Now

Senior Analyst, Product & AI Engineering

Tata Consultancy Services · Fortune 500 financial services client · Plano, TX

  • Owned roadmap and PRDs for an internal analytics platform serving Finance, Operations, and Product — cut executive time-to-insight by 30%
  • Defined the evaluation and quality bar for a production multi-agent AI system shipped to 98% automation
  • Led a 4-person data engineering team; drove delivery across engineering, data, and business stakeholders in Agile cycles
  • Built governance and responsible-AI controls for a regulated environment — audit logging, access controls, human-in-the-loop review

2020

Data Analyst, Ad Intelligence Product

Airfind Corp · Remote, US

  • Took product ownership of an early-stage ad analytics product — KPIs, A/B test methodology, behavior dashboards; recommendations drove 15% revenue growth

2017 — 2019

Research Analyst, Automation & Data Quality

Media.Net Software Services · Mumbai, India

  • Built Python and SQL automation pipelines for data quality and analytics workflows — cut processing time 40%

2019 — 2020

MS, Information Technology & Analytics

Rutgers University · Newark, NJ · GPA 3.83

  • Machine learning & statistics, data analysis & visualization, business analytics programming

Skills at a glance

AI Product Management · Product Strategy · PRDs · Roadmapping · User Discovery · LLM Evaluation · LLM-as-judge · Eval Pipeline Design · Agentic Systems · Multi-Agent Orchestration · RAG · Prompt Engineering · Responsible AI · Python · SQL · Snowflake · AWS · dbt · Docker · ETL & Data Pipelines · A/B Testing · KPI Design · Agile / SCRUM · Stakeholder Management

AWS Certified AI Practitioner (AIF-C01) · ActiveAI for Product Management · Pendo