Methodology in action // enterprise validation

Four engagements, measured the way leadership measures them.

Code review bottlenecks, product-to-engineering handoffs, sprint reliability, and enterprise-scale delivery rescue. Each one includes the numbers that held up under formal review — and the ones that couldn't be cleanly measured.

How the work reads

Most of a ticket's life is spent waiting.

An illustration of the pattern the diagnostic looks for, not a result from any single engagement. Every case study below reports its own measured numbers.

Where the time goes
typical
Coding is the short segment
idle
code
review
Upstream specification idle time — unclear intent, incomplete business rules
Active coding time
Downstream verification latency — large PRs sitting in the review queue
What the work targets
levers
The wait-states, not the coding
4.5d
2d
3.5d
Requirements written to an agreed standard before they reach engineering
PR size caps and WIP limits aimed directly at the review queue
Board and status design that makes code flow visible
A measurement baseline, so the next change can be proven
Client details are anonymized by default to protect ongoing relationships. Names and references are available on request.
Case Study 1 of 4

Untangling a Code Review Bottleneck Mid-Cursor-Rollout

A multi-billion dollar premium consumer enterprise executing the largest website and multi-platform mobile architecture redesign in its corporate history, spanning 12 engineering pods, multiple external vendors, and 750 cross-functional stakeholders (REI).

Teams involved

5

Engagement length

6 weeks

The situation
A senior engineering director at a large enterprise SaaS company (details anonymized to protect the relationship — name and reference available on request) brought me in to work with five of his teams, who had just started using Cursor. He'd seen a different kind of process work I'd done elsewhere in the org — this ask was narrower: look specifically at how code was actually moving through the system.
The problem
Everyone had a theory about where time was going. Nobody had data. PR queues were backing up, and there was no shared, factual picture of where the friction actually lived.
What we did
I ran value-stream mapping sessions with most of the developers, then combined what came out of those with their Jira history and raw Git metadata — pulled into Tableau to essentially hand-build the kind of code-flow visibility that a tool like LinearB now offers as a product. That surfaced specific, fixable choke points, so we reworked their Kanban boards and Jira statuses to make code flow visible in a way it hadn't been. We put new team working agreements in place — a cap on PR size, WIP limits — aimed directly at the review queue. We also walked through their last three major releases together as a team, mapping exactly where each one lost time, and ran experiments off what we found. One that stuck: “pair-prompting” — the old pair-programming habit, applied to AI prompting instead. In retro, engineers consistently said it made review feel lighter. We didn't isolate that on its own — it's folded into the velocity number below.
The result
A clean baseline didn't exist going in, so precise before/after numbers are hard to claim with confidence. What we do have: the team's own velocity tracking — which this organization reports quarterly to leadership, so it's not a soft number — showed roughly a 30% improvement, and the stakeholder's own independent read of the team's pace landed in the same 20–30% range.
Empirical Metrics

~30%

velocity improvement

5

teams

6 weeks

engagement length
This project is what sent me looking for better ways to measure AI-era code flow in the first place — the diagnostic process LeanAI runs today grew directly out of the method I built here.
“Ian knows his stuff and is a pleasure to work with. Through his Diagnostic, he identified the bottlenecks in our code flow and helped us systematically remove them. We were very impressed with the impact he had. He has deep experience in building software and the teams respected him immediately. I would highly recommend working with him.”
Justin N. · Senior Director of Engineering, Tableau
Case Study 2 of 4

Fixing the Product-to-Engineering Handoff Before It Got Expensive

A multi-billion dollar premium consumer enterprise executing the largest website and multi-platform mobile architecture redesign in its corporate history, spanning 12 engineering pods, multiple external vendors, and 750 cross-functional stakeholders (REI).

Engagement length

6 weeks

Focus

PM-to-Engineering handoff

The situation
The work above got me thinking about something I hadn't seen written up anywhere yet: vague requirements don't just slow human engineers down — they get worse once an AI agent is the one turning them into code. I started forming a clear opinion on it, and a peer who heard it hired me to test the theory directly. A senior director of product management at a large enterprise software company (details anonymized to protect the relationship — name and reference available on request) brought me into his org to work on exactly this.
The problem
The handoff from product to engineering had the usual issues — vague language, incomplete business rules — the kind of ambiguity a human engineer used to just resolve with a hallway conversation or an educated guess. That workaround doesn't really exist with an AI coding agent. It either builds the wrong thing fast, or someone has to stop and clarify anyway — either way, you've burned time and tokens on something that shouldn't have needed a second pass.
What we did
We didn't touch how the product org did its own strategic work — the scope was narrow and specific: fix the handoff itself, nothing upstream of it. We built working agreements for how requirements got written — no vague language, complete business rules — and started pairing sessions between PMs and engineers to build Jira tickets together, instead of writing them separately and iterating through rounds of clarification afterward.
The result
The bottleneck itself was hard to see in the data going in — teams weren't visibly parking tickets in a “blocked” state to represent the friction, so there wasn't a clean before-picture to measure against. What we could measure: overall cycle time for feature tickets through the system dropped 23–25%. Engineering also reported a noticeable drop in code churn, though we didn't have a clean way to quantify that specifically in this engagement. The one number I most wanted — the actual change in AI token usage before and after — I wasn't able to measure at all here. I still regret that. It was the whole point of my original hypothesis, and I never got the data to prove it directly.
Empirical Metrics

23–25%

cycle time reduction

6 weeks

engagement length
This is the engagement that convinced me product management can't sit outside the AI-coding conversation — it has to be inside it. It's part of why LeanAI's diagnostic spends as much time on the requirements handoff as it does on the code itself.
“Ian leads with integrity. He worked hard to find ways we could make our practices more efficient and his curiosity and desire to find true bottlenecks and improvements helped us tremendously. AI has changed how we work and Ian helped us become faster and smoother in the AI world.”
Joe Mazzotta · Director of Product Management, ServiceNow
Case Study 3 of 4

Turning Missed Sprint Commitments Into a 95% Delivery Rate

A multi-billion dollar premium consumer enterprise executing the largest website and multi-platform mobile architecture redesign in its corporate history, spanning 12 engineering pods, multiple external vendors, and 750 cross-functional stakeholders (REI).

Teams involved

27

Engagement length

~2 years

The situation
I worked with a VP of Engineering at a large enterprise software company (details anonymized to protect the relationship — name and reference available on request) across 27 teams, moving them off Scrum and onto Lean and Kanban.
The problem
Teams were consistently missing what they committed to, at both the sprint level and the quarterly level — the kind of gap that quietly erodes trust between engineering and the rest of the business, no matter how good the underlying work actually is.
What we did
We moved all 27 teams from Scrum to Lean and Kanban — not a light retooling of ceremony names, but a real change in how work was sized, tracked, and committed to. Getting delivery reliability to where it needed to be took about two years, with results reported formally every quarter — I personally defended those numbers to the full VP group each cycle. Once individual teams were performing consistently, I spent a third year building and delivering a workshop — “Best Practices for Delivering Cross-Team Projects in an Agile Organization” — which she made mandatory attendance across her entire org, then spent several quarters shadowing individual project leads directly, helping them apply it to their own cross-team work.
The result
Sprint delivery reliability went from 55% to 95%. Quarterly commitment delivery went from 55% to 85%. Both numbers came from the same formal quarterly review every org at the company went through — not something I calculated myself after the fact. By that measure, this became one of the best-performing organizations within our division.
Empirical Metrics

55% → 95%

sprint delivery

55% → 85%

quarterly delivery

27

teams
This is the model for LeanAI's Flow Engineering, Phase 1 and Phase 2 — real operational discipline, not more ceremony, with numbers that hold up under scrutiny from leadership, not just a team retro.
“Ian's a great partner to help you drive better practice and better results. He combines Lean delivery practice, empathy and a deep understanding of how teams operate to help them accomplish their business goals.”
Ellie Fields · Chief Product & Engineering Officer, Salesloft
Case Study 4 of 4

Rescuing the Largest Project in Company History

A multi-billion dollar premium consumer enterprise executing the largest website and multi-platform mobile architecture redesign in its corporate history, spanning 12 engineering pods, multiple external vendors, and 750 cross-functional stakeholders (REI).

Teams involved

12

Engagement length

18 months

The situation
I was the senior program manager on the complete redesign of a national outdoor retailer's e-commerce site and its three mobile apps (details anonymized to protect the relationship — name and reference available on request) — the largest project in the company's history, touching roughly 750 stakeholders.
The problem
By the time I was asked to take over, the project was three months behind schedule, and stakeholders — including the CEO — were openly frustrated.
What we did
I introduced an enterprise-level Kanban model across 12 teams and 2 external vendors, running it for 18 months. I set up recurring executive updates — roughly every two months — specifically to keep leadership informed rather than surprised, and ran regular retrospectives that included stakeholders directly, not just the delivery teams.
The result
We delivered one month early — a swing of roughly four months from where the project stood when I took it over. At a retrospective with more than 300 people in the room, the team's own read of the approach was overwhelmingly positive: the consensus was that the Kanban model had made the project transparent and had saved it from real trouble.
Empirical Metrics

1 month

ahead of schedule

300+

at the retrospective

12 teams

2 external vendors
This was enterprise-scale coordination — a dozen teams, external vendors, hundreds of stakeholders. The same value-stream and Kanban discipline applies just as directly at the roughly 100-engineer scale LeanAI typically works with — usually more cleanly, with fewer layers and handoffs to cut through.