Today we are publishing The CTO Playbook for Agentic Systems on our white papers shelf. It runs to 71 pages across nine parts, and it was written by Andrew Stevens, our CTO and CISO (Stevens, 2026), over eleven months of conversations with engineering leaders running agents in production. It is published by Whitepaper Press rather than by us. We are hosting it here because its audience is the one we spend most of our time with.
Why now
Most of the engineering organisations we work with have stopped asking whether agents can write production code. The question has moved on to what happens to an organisation once a large share of its output arrives without a human having typed it, and the public telemetry on that has turned uncomfortable.
Faros Research’s 2026 study, drawn from two years of data across 22,000 developers and more than 4,000 teams, reports throughput up on every measure it tracks, and alongside it a 242.7 percent rise in the ratio of production incidents to merged pull requests, bugs per developer up from 9 percent in its 2025 edition to 54 percent, and pull requests merged with no review at all, human or agentic, up 31.3 percent (Faros Research, 2025; Faros Research, 2026). Faros does not read that last figure as anyone deciding to skip oversight. It reads it as reviewers being unable to keep pace.
We would hold those findings a little loosely, since they are telemetry from organisations running one vendor’s platform and DORA’s 2025 survey reaches a more optimistic conclusion about whether strong engineering practice protects you (DORA, 2025). But if your own review load has started behaving anything like that, the shape of the problem is already familiar and the answers are not obvious.
What the paper works through
Nine parts, each closing with a specific deliverable and a set of questions to put in front of your own leadership team.
Part 2 treats verification as a staffed capability rather than an overhead, sizes it, and puts a formula behind the cost per 1,000 governed actions. It also works the case most write-ups avoid, which is what an agent programme looks like from the inside once it has stopped paying for itself, and when the right call is to shrink it.
Part 3 names seven engineering functions that need owners, with a level range and a compensation-band anchor for each, then raises the problem underneath the ladder: agents absorb a good deal of the work junior engineers used to build judgement on.
Parts 4 and 5 cover how to split work between a squad and its agents, and how to make decision rights auditable. Most first attempts collapse into a binary of what an agent may and may not do, which does not survive an incident review. The paper runs two axes instead, and is direct about the tension that creates between a containment target and a break-glass approval.
Part 6 rebuilds the delivery lifecycle around a different unit of deployment, and includes a sourcing table you can hand to procurement, with a buy-or-build call and an owner for every stage.
Part 7 is the leadership work: five conversations worth preparing for, and a first-90-days communication runbook where every announcement carries a precondition that has to be true before you make it.
Parts 8 and 9 charter a governance board, put fourteen metrics on one page for it, set out four layers of hard control, and sequence the whole thing across six phases with a gate before scale. The gate has five criteria and runs in both directions, which is the design decision we find most useful in it.
What you can actually use
Two things make this more than a read.
The nine deliverables. Every part closes with a named artefact rather than a recommendation, each with an owner, a cadence, and a description of what finished looks like. An agent inventory with named owners. A rehearsed and timed rollback runbook. A signed, dated declaration of which roadmap phase you are actually in and which gate criterion you currently fail. When we are asked to assess an agent programme, these are reliably the documents that turn out not to exist.
The readiness assessment. Twenty-eight questions in the appendix, scored 1 to 5, describing what a weak answer reveals rather than what a strong one sounds like. It is built to be scored in a board meeting. We expect to be asked to run it, and we would rather organisations ran it themselves first.
Who it is for
CTOs and VPs of Engineering with agents already shipping code, the directors and principal engineers designing the review and levelling systems underneath that, security architects who own the control architecture behind an autonomy tier, and the executives who hold engineering accountable.
Download the playbook. For the architecture it assumes underneath, The Trustworthy Agentic AI Blueprint is the deeper treatment and GATE is its implementable form. GATE is an open framework Andrew Stevens authors and maintains personally, not a Sakura Sky product.
If you would rather not work through the twenty-eight questions alone, that is what our Managed GRC service line is for, and a conversation is usually the quickest way to establish which of the nine artefacts you are missing.
Disclosure: The CTO Playbook for Agentic Systems is written by Andrew Stevens, CTO and CISO at Sakura Sky, and published by Whitepaper Press, with copyright held by the author. The white paper is a registration download and Sakura Sky receives the registration details. GATE is an open framework authored and maintained by Andrew Stevens personally under CC BY 4.0 and MIT, and is not a Sakura Sky product. The Trustworthy Agentic AI Blueprint linked here is likewise Sakura Sky and Stevens work, and the same weighting applies. Managed GRC Services is a Sakura Sky offering, and Sakura Sky provides advisory and managed services of the kind discussed here. Third-party research is cited as published and was checked on the dates given in the references.
Not legal advice. This article offers general commentary for an engineering leadership audience. The control-to-framework mapping described in the paper is directional and does not imply conformance with any regime or standard. Readers must obtain independent advice on how any of these apply to their circumstances.
References
DORA, 2025. State of AI-assisted Software Development. Google Cloud. Available at: https://dora.dev/research/2025/dora-report/ [Accessed 18 September 2026].
Faros Research, 2025. The AI Productivity Paradox Report 2025. 23 July. Faros AI. Available at: https://www.faros.ai/blog/ai-software-engineering [Accessed 18 September 2026].
Faros Research, 2026. Ten takeaways from the AI Engineering Report 2026: The Acceleration Whiplash. 12 April. Faros AI. Available at: https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways [Accessed 18 September 2026].
Stevens, A., 2026. The CTO Playbook for Agentic Systems, Version 1.2. Whitepaper Press. Available at: https://www.sakurasky.com/white-papers/cto-playbook-for-agentic-systems/ [Accessed 18 September 2026].

