IB-DD MVP sprint — engineer critique pushback

Branch renamed to ryan-choi-jh/mvp-sprint-plan. I pulled all 48 messages from the claudebin thread and captured the sprint context as memory so it's loaded into every future session in this project.

What's now in my head for the sprint:

  1. Timeline: M1 Apr 24 (workflow + dummy data) → M2 Apr 26 (v1 functional) → M3 Apr 30 (v2 = 80% recall, 0 false positives, dedup layer) → M4 May 1–10 (polish + 5 banker user tests).

  2. The moat: canonical Q/A storage + cross-buyer semantic dedup (tasks 3.5–3.8). Demo headline = "12/20 questions matched buyer 1 — 60% of work eliminated."

  3. Locked rules: zero false positives, conservative classifier bias, Excel-only v1, no OneDrive in MVP, Azure-native stack, pgvector on Postgres, 7 architecture decisions to lock in eng review.

  4. In progress: task 1.1 — user flow storyboard, full May 10 scope, FigJam recommended.

Where you left off: about to start on 1.1 (user flow storyboard). The pre-1.1 checklist I gave you was: confirm full-demo scope (✅), pick a medium (FigJam recommended), write a happy-path narrative paragraph in plain English before drawing.

Want me to (a) help you write that happy-path paragraph now so you can FigJam against it, (b) start on a different task, or (c) something else?


something else before jumping in — i just had a call with an engineer and what he essentially said was:

  1. the product i'm trying to build seem to ultimately just boil down to context management because if you could dump everything into chatgpt and get an answer, then that's about managing context

  2. is domain-specific knowledge really important? otherwise #1 is essentially the problem that everyone is trying to solve

  3. to test whether this is something bankers would actually want, you don't even need an excel-like output interface, but just need a chatbot that can ingest all the data and tell you what you need to know

thoughts on this? i agree to some degree but i feel like it


The engineer is right on two of three points at the abstract level, but each one collapses under domain pressure. Let me take them in order from first principles.

Point 1: "It all reduces to context management"

True at the abstract level. Strategically meaningless.

By that logic Cursor is "just context management for code," Harvey is "just context management for legal docs," Hebbia is "just context management for finance docs." Each of those is a $1B+ company. The reductive frame is the same one VCs used in 2023 to dismiss every vertical AI play — and it kept being wrong.

The actual question is: where does context management cross the threshold from "you could do it in ChatGPT" to "no banker actually does"? Three thresholds:

  1. Corpus size. A typical M&A VDR is several thousand documents and tens of thousands of pages. Even Gemini's 2M-token window can't hold it, and the "dump everything in" approach degrades sharply once you cross ~200K tokens of mixed PDF noise. Retrieval isn't an optimization — it's required.

  2. Throughput. Buyer requests are 100–500 line items across 10–30 buyer groups per deal. ChatGPT is an interactive interface; this is a batch problem.

  3. Trust artifacts. Banker can't send a chatbot transcript to a buyer. They need cited excerpts, file references, audit trail, structured columns. ChatGPT doesn't produce that natively.

The engineer's framing is true the way "all software is just moving bytes" is true.

Point 2: "Is domain knowledge really important?"

This is his sharpest point, and the place where most vertical AI startups DO die. But he's locating "domain knowledge" in the wrong place.

The model isn't where domain knowledge lives in a winning vertical AI product. It lives in three other places, and frontier LLMs don't help you with any of them:

  1. Workflow knowledge. What bankers do, in what order, with what artifacts. The 21 discovery interviews you ran ARE the moat — that knowledge is encoded in scope decisions, not weights.

  2. Eval harness. Harvey's edge isn't prompts; it's their labeled eval set across thousands of legal tasks. Your task 3.1 (golden dataset for diligence classification) is the same kind of asset. ChatGPT can't build it for you because the ground truth lives in banker heads.

  3. Trust UX. Conservative bias, "Partial" as a first-class output, source-passage verification, dedup confirmation flows. These are product decisions made because you understand bankers won't tolerate false positives. A general chatbot makes the opposite tradeoff (it'll always answer, never abstain) because it's optimized for consumer engagement, not regulated workflow.

Data point: OpenAI launched GPT-4 with vision and large context in 2023. Harvey's revenue grew through 2024 and 2025 anyway. If general LLMs were eating vertical AI, Harvey's chart would look different. Same story for Hebbia in finance, Cursor in code, Abridge in healthcare.

The pattern is consistent: general models get smarter; vertical workflow products keep winning where workflow + trust + integration matter more than raw IQ.

Point 3: "Just build a chatbot to test demand"

This is where his critique is partially right and you should take it seriously.

He's wrong that a chatbot tests the product. The pitch isn't "ask the VDR questions one at a time" — that's what bankers already do via email with juniors. The pitch is "turn a 2-week batch workflow into a 2-hour batch workflow." A chatbot reproduces the current pain at lower friction, which is interesting but not transformative.

He's also wrong that a chatbot tests the dedup compounding story — that requires structured Q/A storage and batch processing. It can't surface in chat.

But he's right about the underlying instinct: you should be able to prove demand before sinking 2 weeks into the build. And there's a faster test than either the chatbot or the full sprint:

Concierge-mode validation (1–2 days, before May 10):

  1. Get one friend banker to send you an anonymized real buyer request + access to a sample deal folder.

  2. You + Claude/ChatGPT manually process it. Hand them back the enriched tracker. Time it.

  3. Watch their reaction. If they say "I'd pay for this" → sprint is right. If they say "interesting, but I'd just do it myself" → you have data that should reframe the sprint before you spend it.

This is the steel-man of his point: prove the workflow saves enough time to matter, before building the workflow.

Net read

The engineer is making a generic AI-skeptic argument that doesn't survive the specifics of M&A diligence. But he's surfaced one genuinely useful instinct — validate before building — that your sprint plan currently skips. You jump from discovery interviews (validated demand for the problem) straight to building (assumes demand for your specific solution). Concierge-mode test in the next 2 days closes that gap without slowing the sprint meaningfully.