Skip to content
Article

Governing AI Planning and Code Execution With Branches, Draft Merge Requests, and Audit Records

AI agents that plan work and propose code need the same gates content pipelines do. Here is how we encode governance in branches, draft merge requests, and audit records—while humans keep merge and publish.

Give an agent a planning loop and a code editor, and it will change your repository. That is useful until nobody can say which change was intentional, who authorized it, or whether main was protected when the model got clever. The hard product question is not “can AI write a patch?” It is “can AI plan and propose code while leaving a reviewable artifact trail that humans still control?”

We ask that question the same way we asked it for content. In our earlier essay on an autonomous content pipeline with a human publication gate, the rule was simple: agents may discover, summarize, and draft; humans publish. This piece is the sibling for planning and code. Agents may branch, open draft merge requests, and write audit records. Humans keep the merge gate and the publish gate. Deploy is never an autonomous claim.

TL;DR

Governing AI that plans work and proposes code means encoding safety in the same places your team already trusts: feature branches so main stays protected, draft merge requests so proposals get CI and human review before they become ready to merge, and audit records so every MCP call, evidence packet, and draft status is attributable. Humans approve merge and publish. Agents do not “ship.”

The problem chat-only agent runs create

Chat-only agent sessions feel fast because they leave almost nothing behind. A model plans in a thread, edits files on a laptop, maybe pastes a diff into Slack, and the run evaporates. There is no branch name to bisect, no merge request for CI, no record of which tool wrote which file. When something breaks, you are reconstructing intent from memory.

Auto-merge fantasies make it worse. “The agent opened a PR and merged it” sounds like autonomy. In real ops it is a missing gate. Merge and publish are irreversible brand- and production-bearing actions. If your stack cannot say “the agent cannot merge” and “the agent cannot publish,” you do not have governance. You have a hope with a nice transcript.

We treat that as an architectural constraint, not a policy PDF. The same way create_blog_post_draft and publish_blog_post are different calls in Ansible, branch creation and merge are different privileges in Git. Draft MR is not merge. Draft CMS post is not live. Soft-fail enrichment (hero images, taxonomy guesses) must never skip the gate.

Pattern A — Branches as the unit of agent work

Agents that change code should work on feature branches. Main (or your protected default) stays human-gated. That sounds obvious until you watch a harness default to committing on the default branch because “it was faster.”

Branch discipline gives you three governance properties for free:

Isolation. Agent experiments cannot silently rewrite the shared tip. CI and review attach to a named line of history.

Attribution. Branch naming and commit authorship should make agent vs human obvious—prefixes, bot identities, or MCP actor metadata—so later archaeology is not guesswork.

Reversibility. A bad agent branch is deleted or never merged. A bad commit on main is an incident.

In practice we expect Cursor cloud agents and similar harnesses to open a branch (or work on an assigned one), push commits there, and stop short of merging. Local “fix it on my machine” runs that never leave a branch are fine for exploration; they are not the production pattern for changes that might ship.

Pattern B — Draft merge requests are proposals, not merges

Once an agent has a branch with a coherent change set, the next artifact is a draft merge request (or draft pull request). Draft is the point. The agent proposes. CI runs. Reviewers comment. A human marks ready, approves, and merges—or closes it.

Draft MRs encode several rules we refuse to weaken:

Evidence in the description. What was planned, what files changed, what tests were run, what confidential names were scrubbed. An empty “agent PR” with no evidence is not a proposal; it is a dare.

CI before human attention. Draft does not mean “skip pipelines.” It means “do not merge yet.” Green checks reduce the cost of the human gate; they do not replace it.

Ready is a human verb. Marking a draft ready to merge, approving, and merging are the publication-equivalent gates for code. Agents may open drafts. They do not mark ready or merge as an unsupervised habit.

This is the code twin of draft-first CMS tools. A draft MR that sits unmerged is success for the automation boundary. An agent that merges because the model “felt confident” is a governance failure dressed as velocity.

Pattern C — Audit records for planning and code proposals

Branches and draft MRs cover the git surface. Planning and tool use need a third layer: audit records.

When agents call MCP tools against Ansible, GitLab, GitHub, or Drive, every mutating action should be attributable: which principal authenticated, which tool ran, what entity was created or updated, and where the evidence lives. We keep evidence packets in Google Drive (validation Docs tied to a draft id, slug, taxonomy, and CTA set). CMS draft status itself is an audit fact—status=draft, pinned=false is a machine-checkable claim, not a Slack vibe.

Audit records answer questions chat logs alone cannot:

  • Who (employee OAuth vs system API key) created this draft or opened this MR?

  • What planning brief or ticket keyed the change?

  • Did the agent claim soft-fail enrichment (hero, taxonomy) or skip it?

  • Was publish or merge ever invoked—and by whom?

Without that trail, “the agent did something” becomes unaccountable change. With it, humans can review planning quality the same way they review a diff.

How this maps to Endertech practice

Inside Ansible (our Symfony CMS and operations platform) and the surrounding agent mesh, the pattern shows up as concrete tool shapes and habits.

Content drafts. Authority and authentic posts land via create_blog_post_draft or the GitLab discover → summarize → generate engine. Automation forces draft. publish_blog_post and authentic pin decisions stay human. Optional generateHeroImage soft-fails so missing art never becomes an excuse to skip the gate.

Code proposals. Cursor cloud agents and similar runners open branches and draft PRs/MRs against GitHub or GitLab. Humans review, mark ready, and merge. We do not describe deploy or merge as autonomous, even when CI is green.

Evidence packets. Validation Docs in a shared Drive folder record purpose, CMS ids, confidentiality checks, and next human steps (voice pass, publish approval). That is the audit sibling of the MR description.

Same philosophy, different surface. The content-pipeline essay was about discover → summarize → generate with a human publication gate. This essay is about branches → draft MRs → audit records with a human merge gate. Both refuse the fantasy that agents should own irreversible brand- or production-bearing actions.

Client proof stays anonymized: craft lessons and percentages, never gated brand names. Pipelines that can read private repositories make that rule non-negotiable at the human gate.

A practical checklist you can copy

  1. Require feature branches for agent code changes; protect main.

  2. Name and attribute branches/commits so agent vs human is obvious later.

  3. Have agents open draft merge requests with evidence in the description—not silent pushes to main.

  4. Run CI on drafts; treat green checks as input to humans, not permission to auto-merge.

  5. Keep mark-ready, approve, and merge as human actions.

  6. Keep publish (and authentic pin) as human actions on brand surfaces.

  7. Log MCP/tool attribution and store evidence packets (Drive Docs, ticket comments, MR notes) for planning and proposals.

  8. Soft-fail optional enrichment so gates stay about judgment, not waiting on an image API.

  9. Never claim agents deploy or merge autonomously—even in marketing copy.

  10. If your stack cannot say “the agent cannot merge” and “the agent cannot publish,” fix the tools before you scale the agents.

Governance is not a personality trait of a careful model. It is whether your workflow artifacts make the unsafe path hard and the reviewable path default.

Where this goes next

We are still tightening how planning briefs, cloud-agent branches, draft MRs, and CMS drafts share one vocabulary of draft-versus-irreversible. The interesting work is no longer “can an agent open a PR?” It is whether an operating system for agents can keep humans decisive at merge and publish while agents keep the queue full of grounded proposals.

If you are designing the same boundary—or want help standing up an MCP-backed agent mesh around your CMS and git hosts—start with the gates, not the model name.

Explore how we approach this in practice on our AI agent orchestration solution page, or talk with us about an AI workflow for your team.

Drag to pan. Use +/− or Ctrl/Cmd + scroll to zoom. Pinch to zoom on touch devices.