Seven Questions Every AI Vendor Should Be Able to Answer
The operational diagnostic for vendor-trust collapse risk
§1. Why this matters
The previous piece in this series, Cognitive Surrender, named the failure mode and catalogued seven flavours: AI-builder pricing, GTM-tooling pricing, platform enforcement, agency AI substitution, the boss-obsessed-with-AI pattern, infrastructure enshittification, and the seven-figure-ChatGPT-wrapper transformation theatre. The taxonomy is useful, but on its own it is inert.
This piece is the operational complement. If you are the CIO, CISO, risk-committee chair, GC, or procurement officer approving an AI vendor next quarter, you do not need another thesis. You need a diagnostic you can run in a discovery call, an RFP, or a 30-minute vendor-eval review. Seven questions, one per flavour, with vendor-honest and vendor-evasion answer shapes laid out so you can tell which one you are hearing.
Vendor-honest answers are short, specific, verifiable. Evasions are abstract and conditional. You do not need to be the most technical person in the room to tell them apart; you need to know what to listen for. Each question below maps to one flavour, with the honest answer shape, the evasion shape, follow-ups, and a receipt from the public record.
§2. Question 1 — Is your pricing tied to volume, or to correct outcomes?
This is the AI-builder pricing question fused with the transformation-theatre question. Both flavours share an economic structure: the vendor profits when the customer stops checking.
The vendor-honest answer pairs pricing with correctness: refund-on-wrong, outcome-tied milestones, deterministic-recipe unit pricing, or some combination. The vendor names the correctness metric, the threshold below which the customer pays nothing, and the audit method that establishes whether the threshold was met. This does not require the vendor to be cheap; it requires the vendor to put margin at risk when the work is wrong.
The vendor-evasion answer is abstract: “Our pricing reflects value delivered.” “Credits give you flexibility.” “Per-action metering scales with usage.” None ties payment to correctness; each transfers wrongness risk to the buyer. The most candid version of the evasion was posted to r/replit by Michele Catasta, who runs Replit as President under the handle pirroh: “Predicting the estimated cost is a very hard technical problem. We don’t do it today not because we don’t want to, but because the estimates would be wildly inaccurate.” The President of the largest agentic-code platform put in writing that the pricing model is unpredictable because predicting it correctly would expose how inaccurate the underlying work is.
Follow-ups when the vendor evades:
What percentage of last quarter’s revenue went back as refunds for incorrect outputs? If zero, the refund policy is theoretical.
Show me a customer invoice from the last 90 days where you took a haircut because the output was wrong.
What is the variance between your billing forecast and actual on a typical engagement?
If we run the same workflow ten times with identical inputs, what is the cost variance across those ten runs?
A vendor whose pricing is tied to correctness answers all four with specifics. A vendor whose pricing is tied to volume redirects to “value” or “flexibility.”
§3. Question 2 — Where is the audit trail, and what does it record?
This is the GTM-tooling question fused with the transformation-theatre question. Audit trail is the artifact that proves the work survives scrutiny. Its presence or absence tells you whether the work was structured to be checked.
The vendor-honest answer names the unit of record: every action, model call, tool invocation, approval gate, input prompt, retrieved context, and output, with tenant attribution, timestamps, the reasoning chain, and the approver identity if the action was gated. The trail is replayable: you can pull a decision from 90 days ago and reconstruct what the system knew, what it concluded, who signed off, and which deterministic recipe ran. Retention is customer-controlled. The format is exportable.
The vendor-evasion answer is shallow: “We have logging.” “Audit is on the roadmap.” “You can see your action history in the dashboard.” The r/sysadmin transformation-theatre thread captured what evasion-grade logging looks like in production. bigbadrune on r/sysadmin (⬆2,326): “our ‘ai transformation’ cost seven figures and delivered a chatgpt wrapper... literally a system prompt that says ‘you are a helpful assistant for [company name]’. same hallucinations, same limitations, except now it confidently makes up internal policies that don’t exist.” ruibranco replied: “’full admin permissions to anything’ is the part that would keep me up at night.” The wrapper logged that it answered. It did not log what was wrong with the answer.
Follow-ups when the vendor evades:
Replay one decision from 90 days ago end to end, including retrieved context, reasoning chain, output, and approver identity.
Does the trail include the prompt and context, or only the final action? Action-only logs will not catch a hallucinated input.
Is the approver identity logged as part of the immutable record, or only the action they approved?
Can my compliance team extract the audit trail in a structured format, on demand, without your cooperation?
If the vendor cannot replay a decision and reproduce its reasoning chain, the audit trail is decorative. The artifact’s job is to survive scrutiny six months later, not to populate a dashboard widget.
§4. Question 3 — When did your platform last get banned, suspended, or rate-limited by a partner?
This is the platform enforcement question. Almost no vendor will answer it cleanly, because the honest answer is “more often than we tell prospects.”
The vendor-honest answer is full disclosure: the event, root cause, mitigation, contingency. “In Q3 we hit a rate-limit threshold with [partner] that affected 14% of our customer base for 48 hours. Root cause: an internal policy change at the partner. Mitigation: reduced throughput, reworked the integration pattern.”
The vendor-evasion answer is one of two reflexes: “we comply with all platform policies” or “we have not experienced any service disruption.” Both are technically true and operationally useless. They tell you nothing about what happens when the platform changes its mind. The receipt is LinkedIn enforcement: a 40% year-on-year ban rate of HeyReach, Apollo, Lemlist, La Growth Machine, and Amplemarket profiles, surfaced by Material_Hospital_68 on r/GTMbuilders this year. Forty per cent of the install base of the dominant outbound-automation stack was banned by the platform it depends on, over twelve months. Disclosure is voluntary, and the incentive to disclose is zero.
Follow-ups when the vendor evades:
Has any partner platform suspended, banned, rate-limited, or formally warned your service or your customers’ accounts in the last 24 months? Yes or no.
If yes: name the partner, the date, the customer-impact percentage, and the disclosure you made to existing customers at the time.
What is your contingency if LinkedIn, Slack, Microsoft, Google, or any major partner bans your integration tomorrow? Walk me through it.
What is your portability story? If a partner change makes you non-viable, what does our migration off you look like, and how long does it take?
A vendor with a credible platform story answers all four with specifics. One whose business depends on tolerated grey-area access redirects to compliance certificates.
§5. Question 4 — Are you the substitute for an employee, or the supplement to one?
This is the agency-AI-substitution question fused with the boss-obsessed-with-AI question. Both are really one question: what is the human’s role in the workflow you are selling?
The vendor-honest answer is “supplement”, and the pricing reflects it: per-active-user pricing, scoped task definitions, named approver gates, judgment-loop preserved where judgment is required. The vendor articulates which decisions humans own and the failure modes when humans are removed.
The vendor-evasion answer reaches for substitution language: “AI worker”, “10 AI employees for the price of 1”, “replace your sales team”, “fully autonomous”. The receipts are extensive. r/marketing #24 (⬆426, 223 comments). asp821: “I have a client like this. ChatGPTs everything to death until there’s nothing memorable about it. Sunday we launched an ad that’s done better than any other ads we’ve done in awhile and he immediately went in there and started changing shit after running it through ChatGPT.” Logical_Bite3221: “Boomer, obsessed with AI... wants us to use it for everything so we are teaching it how to do our jobs so we can be laid off soon.” r/content_marketing carries the agency-layer parallels: the lead designer who quit when an agency forced AI creative on a premium client (#5), the 12-year writer whose career feels over (#13), the 5-person ad agency that laid off 2 of 5 because only one person knew the AI workflow (#6).
What the substitution claim hides is the rework: human time fixing AI output, client trust rebuilt after an AI artefact slipped through QC, layoff cost when substitution proves uneven. The honest vendor prices to the rework. The evasive vendor prices as if rework will not exist.
Follow-ups when the vendor evades:
Walk me through the workflow end to end. At each step, name which decisions humans own. If everything is model-owned, where do errors get caught?
What is the average human-time per AI output in production? “Near zero” means either unmeasured or overstated.
What is the rework policy when AI output requires substantial revision? Define “substantial.”
Show me three customers running the workflow at six months or more, with human-headcount and per-output review time before and after.
If the answer to follow-up 1 is “the model owns everything”, the vendor’s product depends on you surrendering the judgment loop. The alternative is paying slightly more for a vendor whose pricing treats the human as a contracted gate, not a friction point.
§6. Question 5 — What happens when our incumbent vendor enshittifies?
This is the infrastructure enshittification question. It is rarely asked at selection because it sounds defeatist, and it is exactly the question that protects the buyer five years into the relationship.
The vendor-honest answer is a portability story. Open APIs, customer-controlled data export in a documented format, BYOK for any credentials the vendor uses to act on the customer’s behalf, contractually defined termination notice with migration assistance, no proprietary lock-in on the workflow definition. The vendor articulates what the buyer takes on exit: custom workflows, training data, audit trails, integration configurations.
The vendor-evasion answer is soft assurance: “we’d never do that to you.” Trust is not a portability story. The receipts are ones every IT operator already knows. r/sysadmin’s top post of the year (⬆9,196) carried the Broadcom-VMware playbook: “Your licenses expire today and you will face environment disruptions as well as penalty fees” against perpetual licenses, captured by MeridianNL as “you got upgraded from ‘customer’ to ‘hostage’.” The same thread surfaced GitHub Actions self-hosted runner fees adding “close to $3.5k a month extra” for runners the customer already owns. r/sysadmin #19 on Atlassian Rovo bundling, Just_the_nicest_guy: “the enshittification will continue until profits improve.”
Every one of those vendors looked safe at acquisition. Buyers with portability left. The ones without it absorbed the price change.
Follow-ups when the vendor evades:
What is the file format on export, and can I receive a full export on demand?
What is your termination notice period, and what migration assistance is contractually defined during it?
If you pivot or are acquired, what happens to my custom workflows? Are they portable to a competing vendor, or expressed in a proprietary DSL that locks me in?
What are your published price-change policies, and what is the contractual cap on annual price increases?
A vendor with a credible enshittification story makes the buyer’s exit cheap by construction. One that handwaves it is asking for trust on a topic where every major incumbent has demonstrated trust is misplaced.
§7. Question 6 — Show me one real customer in production with you for 18 months or more
This is the transformation-theatre question, asked sideways. The MIT 95%-fail study on enterprise AI pilots gave the operator tier the data point it was reaching for: most pilots never reach production, and vendors with no production customers at the 18-month mark have not been tested by deployment reality.
The vendor-honest answer is a named customer with specifics: deployment date, production-load numbers, the technical contact who will take a call, the list of things that have broken since deployment and what was done, and churn data if other early customers left. Honesty about scale: “this customer runs 30,000 monthly transactions; we have smaller ones, but this is our most production-mature reference.”
The vendor-evasion answer takes one of four shapes: “we just launched”; “our customers are under NDA”; “we’re in stealth with our enterprise references”; “we have a Fortune 500 customer but I can’t say who.” None tells you whether the product survives production; each transfers validation risk to the buyer. The receipts are the public failures. Delve, YC W24 cohort with a $32M Series A, was caught issuing SOC 2 and ISO 27001 audit reports to 494 companies, with 99.8% template-identical and audit conclusions written before the audits started (r/startups #3, ⬆822, 194 comments). Tibo, the solo-builder case on r/lovable #27: a vibe-coded application with 6,927 paid users shipped with an admin-access exploit giving any logged-in user full access to sensitive data across the user base. Both would have had glowing references at month three. Neither would have survived to month eighteen.
Follow-ups when the vendor evades:
Give me the production-load numbers from your most mature deployment: monthly transactions, peak concurrency, data volume, integration count.
What is the technical contact, and will they take a 20-minute call with me before I sign?
What is the deployment date, and what specifically has broken since the customer went live? Top three incidents.
What is your churn on customers acquired more than 12 months ago? How many are still active, how many migrated off?
A vendor with a real 18-month production reference answers all four unprompted; the reference is the easiest sale they can make. A vendor without one redirects to roadmap, vision, and stealth-customer claims.
§8. Question 7 — When we audit you in six months, what evidence will you provide?
This is the meta-question. It crosses every flavour. The audit is not a one-time gate at procurement; it is the ongoing assurance the vendor does not drift toward enabling cognitive surrender as the relationship matures. The question forces the vendor to describe the artifact they will hand over, not the certificate they will display.
The vendor-honest answer is an artifact list. Every customer action logged with full context. Every approval gate logged with approver identity and timestamp. Replayable decisions, customer-controlled retention, structured export. Vendor-side audit data showing what the vendor itself did with the customer’s data, including access by employees and subprocessors. The vendor names the format, the access pattern, and the verification method. The customer extracts the evidence without vendor cooperation.
The vendor-evasion answer is the certificate reflex: “SOC 2 Type II.” “ISO 27001 certified.” “We renew our compliance attestations annually.” r/ITManagers #30 surfaced the buyer-class vocabulary. Necessary_Durian_327: “they all want to know how your ask aligns to the business strategy... What is the Total Cost of Ownership?” r/ITManagers #25 packaged the MIT 95%-fail finding as a procurement blueprint: “partner don’t build / verticals not horizontals / future-proof and adaptable.” The shift is evidence, not opinion. A SOC 2 certificate is opinion at a point in time. A replayable audit trail of every decision the vendor made on the buyer’s data is evidence. Only one survives a six-month review when the question is “what did this vendor actually do with our data, and was it correct?”
Follow-ups when the vendor evades:
Who has access to the audit logs: the customer, the vendor, or both? Vendor-only access means the audit is not customer-verifiable.
What is the retention period, and is it customer-controlled? Vendor-controlled retention can be quietly truncated before it is needed.
What is the export format, and can I extract the trail without vendor cooperation? If cooperation is required, the audit is not independently verifiable.
What vendor-side audit data is available? What did your engineers do with our data, what did your subprocessors do, and where can I see that activity logged?
A vendor whose answer is “we’ll provide what’s needed” is offering the certificate reflex. One who names the artifact, the format, and the customer-controlled access pattern is offering the discipline you are paying for.
§9. How to use these seven questions
These are not gotcha questions. They are the buyer keeping the judgment loop alive at the point where it most often gets surrendered: the AI vendor evaluation, where the seller’s incentive is to compress diligence into a slide deck and the buyer’s is to close the procurement cycle.
A sequencing that has worked in field conversations.
In initial discovery, ask Q1 (pricing), Q2 (audit trail), and Q5 (enshittification). The first two tell you whether the vendor’s economics are aligned with your correctness. The third tells you whether your exit is engineered or improvised.
In deeper evaluation, ask Q3 (platform enforcement), Q4 (substitute or supplement), and Q6 (18-month production reference). These surface whether the vendor has been tested by reality, by partners, and by your workflow shape.
In contract negotiation, ask Q7 (audit artifact at six months). This decides whether the discipline survives renewal. Bake the answer into the contract.
Use these in any RFP, discovery call, or vendor-eval review. No JieGou attribution required. The operator tier needs a clean diagnostic vocabulary; the seven questions exist to be screenshot-shared, RFP-quoted, and put in front of the next AI vendor you meet. The framework in circulation matters more than the credit.
JieGou is structurally built to give vendor-honest answers to each of the seven: deterministic-recipe pricing with named-approver-gate correctness ties, replayable audit trail with customer-controlled retention, no platform-enforcement dependency in the core workflow, named human approvers as structure, customer-controlled data export, production references with technical contacts, and a contract-specified six-month audit artifact list. Run the questions on us, then on every other vendor in your AI stack. Pick the path whose receipts you can live with.
