On a team running coding agents, forty pull requests can land before lunch, and only one of them touches payments. Finding it falls to a senior engineer, who is now the slowest part of the delivery pipeline. Google already generates three-quarters of its new code with AI, a share Sundar Pichai disclosed in April, and engineers still approve it. Every sprint ritual and RFC template in that organization was built to protect engineering hours, and none of them tell the reviewer which of the forty to open first.

Navneet Rao is the Director of Engineering at Thumbtack, where he leads the team behind the home services marketplace's core hiring experience. He shaped Thumbtack's generative AI strategy and drove its adoption across the company. Before Thumbtack, he led machine learning work on IBM's Watson Assistant. His solution is to score each change by its risk and spend human attention where a mistake would be expensive or hard to undo.

"For most of software's history, engineering capacity was the constraint, and organizations built processes to ration it: prioritization rituals, weekly standups, and detailed RFCs before starting major projects. As implementation gets cheaper, judgment, attention, and verification become scarce resources, and those processes should now change," Rao says. AI keeps making developer hours cheaper, and review capacity now faces the pressure that implementation capacity used to. Senior engineers end up drowning in review, and faster code generation has yet to make most delivery pipelines faster overall.

Risk scoring decides agent autonomy

The tempting fix is to slow the agents down, which gives back the speed teams adopted them for. Rao leaves the agents running and spends reviewer attention according to risk.

"Organizations should invest in systems that can measure the level of risk associated with code changes: blast radius, reversibility, sensitivity such as auth, payments, and personal data, and test and eval coverage. Those assessments can then determine how much autonomy AI agents get," Rao explains. The framework turns autonomy into an allocation decision. An agent gets as much independence as the change in front of it can safely bear, so a well-tested edit to an internal tool that rolls back in minutes can merge on its own while a database migration or a change to an authorization system waits for a person. Agent autonomy gets settled one change at a time, with the risk score setting the terms.

"Engineers can shift from reviewing every line of code to focusing on the changes where judgment matters most," he says. The most valuable reviewer on the team becomes the one who knows which changes can hurt the business. Risk scoring makes sure that person's limited hours go to those changes and nowhere else.

Verification still has to scale

The verification problem gets bigger when AI expands who is capable of shipping software. Product managers and designers can now stand up working prototypes, and engineers can cross into iOS or Android work they were never trained for. Rao treats that spread of agency as something to fund, with training on the code quality expectations of unfamiliar platforms and sandboxed environments where teams can try new tools without touching anything live.

"If a PM or designer builds a working prototype, it should be encouraged because it's a faster way to explore new ideas. But as leaders empower people, they need to create systems of accountability. Whoever ships a change should own the outcome, and anything that reaches production or touches customer data should go through rigorous verification," Rao says.

A model-backed feature can pass every conventional test and still start completing fewer tasks, or costing more per request, after a model update. Traditional verification asks whether the code works, and model-backed features add a second question about whether the model is still behaving acceptably weeks after launch. Rao's answer pairs eval datasets that score task success on sample inputs with production monitoring of reliability and cost per task, so a model change that degrades quality gets caught before customers notice. The evals double as a ship threshold for the PMs and designers now building alongside engineers.

"Guardrails can help when they're built into the system, like sandboxes and eval standards, and hinder when they become manual processes that people must do before they can launch experiments," he adds. Approval forms and sign-off meetings scale with headcount, while checks built into tooling scale with the code itself.

A failed experiment can expire

An AI experiment that fails today may only be failing for now. Each model release can remove the limitation that killed a workflow, whether a context window that ran out or a schema the model kept misreading, and the idea only resurfaces if someone wrote down why it broke. "As new models become available, a product idea or workflow that failed a few months back could suddenly become viable. That's why it's important to share successes and also document the interesting failures," Rao says.

That record needs somewhere to live, so he wants teams demoing their AI experiments in open forums, with the workflows that hold up packaged into shared libraries and internal tools other teams can pick up. Google's DORA research found that AI amplifies what's already there inside an organization, so a team that records its dead ends gets a head start every time a new model makes one of them viable.

Rao puts that culture on the person at the top of the org chart. "Asking teams to share failures isn't enough. When leaders share their experiments, especially the ones that didn't work, they reinforce a culture that values context sharing over one that only celebrates success," he says. Leaders willing to fail in public give everyone below them permission to log a dead end without treating it as a career risk. Once building is cheap enough to explore before committing, strategy moves closer to the prototype. A leader can test an idea in an afternoon and walk into the planning meeting with working software, before anyone spends weeks writing the document that explains it.

"Strategy discussions increasingly start from working prototypes rather than documents," Rao says. "With agents summarizing incident threads, pull requests and design discussions, leaders spend less time gathering context and more time making decisions."