GitMir IDE
Book a demo
Start free

AUDIT METHODOLOGY

Every number herecan be argued with.

These are the definitions the audit computes against. They are published because a measurement whose method is private cannot be used in anybody's business case — and because several of them make our numbers smaller rather than larger, which is the point.

DEFINITIONS

Change
One request, and every task that grew out of it — the first implementation, the fixes, the repeats after somebody said it misunderstood. They are linked by a single identifier written when the request first becomes work, so a round that started with a person changing their mind counts the same as one that started with a failed check.
First-pass delivery
Starts when the first task of a Change enters active work. Stops the first time any task of that Change reaches verification. It is deliberately not "until it was accepted" — the first reviewable version is the honest end of delivery.
After the first pass
From that moment to the last transition into done that was not followed by a reopening. Everything in between: clarification, another attempt, another review.
Iteration
One return from verification back into active work. The count for a Change is how many times that happened across all its tasks.
Review cycle
One entry into verification. A change verified once has one; a change that came back twice has three.
Late discovery
A task created after the Change had already reached verification once. That is what a dependency, a rule or a decision found too late actually looks like in the queue.
First-pass ratio
Changes that passed verification once and were accepted, divided by the changes that reached verification at all. Work still in flight is not in the denominator — counting it would make the ratio drift with how busy the week was rather than with how the work went.
Window
A change is in the sample if it STARTED inside the period. One that began earlier is excluded whole, not clipped at the boundary — clipping events instead of changes is how a two-day piece of work reports as "0 minutes to first review".
Idle cutoff
A gap longer than the cutoff — four hours by default, switchable to two or eight — is counted in neither clock. It is a blunt instrument and it throws away real work as well as sleep: moving a task through a queue cannot tell an overnight gap from six hours of concentration. What is measured is time between moves, minus long gaps. It is not "clean working time", and the screen says how many stretches and how many hours the cutoff discarded.
Queue time
Time a task sits untouched before anybody starts it is not counted in either clock. That is backlog, not the cost of the change.

EVIDENCE STATES

What each number is allowed to claim.

Three states, never mixed. A figure carries its state everywhere it appears, including in a proposal.

  1. OBSERVEDComputed from recorded transitions. Timestamps, states, iterations, review cycles. Reproducible from the same data by anybody who has it.
  2. CLASSIFIEDWhy an extra round happened. Inferred, shown as inferred, and correctable by the team. An unreviewed classification is not evidence.
  3. ESTIMATEDAnything involving money. It needs an hourly cost the customer supplied, and it inherits every assumption in that number.
  4. VALIDATEDThe part of the classified cost the customer has agreed is real and addressable. This is the only state an offer may be built on.

WHAT IS NOT MEASURED

And will not be.

  • Individual developers. There is no per-person breakdown in the interface, the API or the export, and one is not going to be added.
  • Anything happening outside the queue. A conversation in Slack, a call, somebody turning round in a chair — real, and not observable, so not counted rather than guessed at.
  • Lines of code, commits, or any other proxy for effort.
  • Time to a deploy. The audit ends at acceptance; what happens after release is a different measurement with different causes.

AND THE PARTS WE ARE NOT SURE ABOUT

Where this is still weak.

A methodology page that only lists strengths is marketing. These are the known limits, and they are the reason the audit is free rather than sold.

  • Classification is inference. Whether an extra round came from a missing product rule or from an ordinary defect is a judgement, and the first version of it will be wrong often enough that a person has to be able to correct it.
  • Sample size is not a number we have. Below four changes the screen says the data is thin rather than showing a confident zero — but four is the threshold for displaying anything, not a claim that four changes mean something. How many it actually takes depends on the team, and we have not measured it across enough teams to publish a figure.
  • The idle cutoff cuts both ways. Four hours splits an overnight gap correctly and six hours of uninterrupted work incorrectly — the queue cannot tell them apart. It is switchable, and the screen shows how much it discarded so the size of the error is visible rather than hidden inside the total.
  • Work that never enters the queue is invisible. If a team does half its changes outside GitMir, the audit describes the half it can see and says so.

See what the audit measuresRun open source

NVIDIA Inception Program member

One living model between business and software development.

Business logic for people and AI. Real changes synchronized. Rework measured automatically.

Open source, local-first, and your source stays in your environment.

Open the app
© 2026 GitMir IDE. All rights reserved.Understand · Execute · Verify