Ask a language model to write a proposal and it will happily write one. It will also invent a price, round up a case study, promise a payback period nobody measured, and describe your team's past work in ways your past clients would not recognize. The prose sounds confident. The commitments inside it are unaccountable.
That is the real problem with AI-generated proposals. Writing is the easy part. The hard part is making sure every number came from a rule a person approved, and every claim came from evidence a person curated. This is the third piece in our series on putting AI to work behind human gates. The content pipeline essay covered drafts and publishing. The governance essay covered branches and merge requests. This one covers the documents that carry commercial commitments: proposals.
TL;DR
In the proposal workflow we built inside Ansible, our Symfony-based operations platform, AI does the reasoning and the writing, while application code owns the commitments. Evidence comes from a curated, versioned library of source-backed claims, never from the model's memory or the open web. Prices come from a deterministic calculator and approved package settings, and any dollar amount the model writes on its own is rejected. Payment terms are derived from approved duration by fixed rules. A claim-safety validator flags urgency, guarantees, ROI talk, and leaked internal mechanics before a salesperson reviews the draft. People approve the strategy, the evidence, and the numbers. The model writes around them.
Why proposals are a harder AI problem than blog posts
A weak blog post costs you a little attention. A weak proposal can cost you a client relationship, because a proposal is a promise. Buyers read it to decide whether the engagement is worth the risk. Every sentence about price, timeline, proof, or outcome is something your team will be held to.
Generic AI drafting fails in three predictable ways:
Invented numbers. Models fill gaps with plausible figures: a package price, an hourly rate, a percentage lift, a "typical" ROI. Plausible is not approved.
Promoted evidence. A capability your team has delivered becomes a "proven result." A client-reported outcome becomes a published benchmark. A planned benefit slips from future tense into fact.
Leaked mechanics. Internal estimating details (hours, labor assumptions, allocation math) end up in buyer-facing prose, where they invite line-item haggling instead of a conversation about outcomes.
None of these are fixed by a better prompt alone. Prompts reduce the frequency of mistakes. Code removes whole categories of them.
Principle 1: Evidence comes from an approved library, not the model's memory
Our proposal system keeps Endertech proof as curated records, one versioned JSON file per record, in the application repository. Each record carries its source URL and retrieval date, the capability or engagement it describes, approved wording for internal review, optional client-facing wording, and two fields that matter most: a claim class and an evidence grade.
Claim classes keep different kinds of truth from blurring together:
verified public fact
published measured result
client-reported result
delivered capability
qualitative outcome
intended or projected benefit
general service capability
The rule is simple: a weaker class can never be promoted to a stronger one. Planned benefits stay in future tense. A delivered capability is described as something we have done, not as a measured outcome we can promise again.
Retrieval is deterministic, too. Before the model sees anything, application code scores the library against the deal (industry, engagement type, platform, capability, risk, and outcome signals), shows the scores and reasons, and trims the candidates to a fixed budget. The model's job is narrower and more useful: decide which candidates are genuinely analogous to this buyer, explain why, and reject superficial matches. The salesperson then sees the selected proof with its source links and can uncheck anything that does not belong.
Two boundaries make this hold. Generation and refinement do not use live web search, so client facts must come from the deal inputs and Endertech proof must come from accepted evidence. And when the recommendation step does reach for something outside the library, such as a relevant article on our own site, it is labeled provisional for human review and is never treated as approved proof.
Principle 2: The model proposes structure; code owns the dollars
This is the heart of the design. In our workflow, the recommendation pass captures option structure only: one option or three, what each one delivers, how long it takes, and how the options genuinely differ. It does not capture package prices.
Pricing happens on a separate Estimate step. AI drafts workstreams and hour ranges for each package from the approved structure, using a fixed, coarse estimation scale so nobody is pretending to know a task is 17.5 hours. A deterministic calculator then turns those ranges into cost and a suggested client fee, applying configured settings and rounding rules. There is no model in that math. When a salesperson approves a package variant, the approved fee is written to the proposal as the source of truth.
After the model returns its structured proposal, generation and every refinement path reject model-authored package dollar amounts whenever structured pricing is configured. Approved investments and payment terms are then merged into the document by application code. Where a planning engagement needs to describe a later phase, the model writes the explanation around placeholder tokens for the low and high bounds, and the application substitutes the approved figures. The model can explain a range. It cannot set one.
The internal costing view (workstreams, labor assumptions, allocation) exports separately as an employee-only document and is never mixed into the client PDF. Buyers see what each option produces and what it costs, not how we built the estimate.
Principle 3: Payment terms follow rules, not vibes
Payment schedules are a good example of something that should never be improvised. In our system they are derived from the approved option's duration:
Engagements under about four weeks are invoiced at kickoff.
Engagements of four to six weeks split between kickoff and completion.
Longer engagements use equal monthly installments over the parsed number of months.
If duration cannot be parsed reliably, the proposal uses general monthly wording with no dollar amounts rather than guessing.
That last rule matters as much as the first three. A deterministic system needs a safe fallback for the cases it cannot compute, and the safe fallback is saying less, not making something up.
The same thinking applies to package design. Every option has to be a real choice with the same quality floor. A premium tier is only priced as premium when it contains genuinely unique scope, and prices are never inflated to make a middle option look good. The model cannot add crossed-out prices, fake discounts, or a deliberately weak "decoy" option, because those patterns are prohibited in the prompts and flagged in review.
Principle 4: Claim safety is checked before a human ever reads the draft
A ProposalClaimSafetyValidator runs before a proposal is saved or rendered. It looks for patterns that do not belong in an honest proposal:
fabricated urgency, scarcity, or popularity ("act now," "limited spots," "most popular")
discount or savings framing
guarantees, "risk-free" language, and unsupported superlatives
ROI and payback claims
leakage of internal pricing mechanics such as hourly rates, margins, multipliers, or allocation math
Not every match is a hard stop. Narrative warnings are recorded in an internal audit field and shown to the salesperson for review, because context matters and a sentence like "results are not guaranteed" is a perfectly honest limitation. Unauthorized dollar amounts are different. Those block the save. The asymmetry is deliberate: style judgments go to a human, while numeric commitments are enforced by code.
Each saved proposal also stores an internal audit record with the validator version, when it ran, any warnings, and the provenance of the accepted evidence. That record never appears in the client document, but it means anyone can later answer "where did this claim come from, and was it checked?"
Principle 5: Approved inputs are locked, and changes make drafts stale
Proposals change as conversations evolve, which creates a quieter risk: generating from inputs that are no longer the ones a person approved.
We handle this with canonical signatures. When the salesperson accepts the buyer context, the recommendation, the selected evidence, and the pricing, the system records a signature over those inputs. If any of them change, the final summary becomes stale. If they change while a background generation is running, that attempt is invalidated rather than quietly finishing with old assumptions.
Final Review is mostly deterministic as well. Accepted context, evidence, pricing options, and signatures are assembled by application code. A model is used only to compress very large sets of CRM activity or documents that exceed a size budget. Before generating, the salesperson can preview the exact request the system will send, including the instructions, the inputs, and the expected output schema, without making a model call. No hidden steps.
Reopening an in-progress proposal does not start AI work on its own, either. People explicitly continue, regenerate, or review previous recommendations. Nothing re-runs just because someone clicked back into a deal.
What AI is actually good for here
None of this means the model is a typist. It does the work that benefits most from reasoning and language:
organizing messy deal inputs into a buyer strategy, with unknowns kept as clearly labeled hypotheses rather than invented facts
asking a handful of targeted discovery questions when the answers would change positioning or scope, and letting the salesperson confirm or correct each hypothesis
drafting genuinely different option structures and first-pass workstreams for estimating
selecting analogous proof and explaining why it fits
writing clear, aspiration-led prose that starts with where the buyer wants to go and treats scope as the path, not the headline
The design principle underneath it comes from value-based pricing thinking: buyers judge outcomes, confidence, and the safety of the decision, not a list of tasks. AI is excellent at helping a team articulate that value. It must never be allowed to guess a buyer's willingness to pay, their approval dynamics, or their ROI. Those are researched or confirmed, not inferred.
A checklist you can copy
Keep proof in a curated, versioned library with source URLs, retrieval dates, and claim classes. Never let a weaker claim class be promoted.
Retrieve evidence deterministically and let the model select and explain, not invent.
Turn off live web search for generation. Client facts come from deal inputs; your proof comes from accepted evidence.
Separate option structure from pricing. Let the model propose structure and draft estimates; let a calculator and a person set the numbers.
Reject any dollar amount the model writes when structured pricing exists. Merge approved figures with code.
Derive payment terms from approved duration with explicit rules, and fall back to saying less when a rule cannot apply.
Validate claims before review: urgency, discounts, guarantees, ROI, and leaked internal mechanics.
Keep internal costing documents separate from client documents.
Sign approved inputs and invalidate drafts when they change.
Let people preview exactly what will be sent to the model, and keep sending the proposal a human decision.
Where this goes next
The pattern across this series is the same: AI earns more responsibility when the irreversible parts stay with people and the commitments are enforced by code. For content, that means humans publish. For code, humans merge. For proposals, humans approve the evidence, the price, and the send, while AI does the heavy lifting of understanding the buyer and writing clearly.
If your team is exploring AI inside the tools you already use for sales, estimating, or operations, start by deciding which numbers and claims must never come from a model. Everything else gets easier once that line is in code.
See how we approach this on our AI agent orchestration solution page, or talk with us about an AI workflow for your team.
