---
# === IDENTITY ===
id: business/product-tech/engineering-productivity-benchmarks/2026
canonical_question: "What are current engineering benchmarks — DORA metrics, cycle time, deployment frequency by team size?"
aliases:
  - "DORA metrics benchmarks"
  - "software delivery performance benchmarks"
  - "deployment frequency benchmarks by team size"
  - "engineering velocity benchmarks 2026"
  - "cycle time benchmarks software teams"
  - "change failure rate industry benchmarks"
  - "MTTR benchmarks DevOps"
entity_type: benchmark
domain: business > product-tech > Engineering Productivity Benchmarks
region: global
jurisdiction: global
temporal_scope: 2025-2026

# === VERIFICATION ===
last_verified: 2026-03-10
confidence: 0.85
version: 1.0
first_published: 2026-03-10

# === TEMPORAL VALIDITY ===
temporal_validity:
  status: volatile
  last_breaking_change: "2025 DORA report added rework rate as 5th metric; shifted from elite/high/medium/low clusters to archetype-based classification"
  next_review: 2026-09-06
  change_sensitivity: high
  data_vintage: "Q4 2025"

# === CONSTRAINTS ===
constraints:
  - "DORA benchmarks are self-reported survey data (~5,000 respondents) — actual measured metrics from platforms like LinearB differ significantly from self-assessments"
  - "Team size dramatically affects achievable benchmarks — a 5-person startup and a 500-person enterprise have structurally different deployment patterns; always segment by team size"
  - "Primarily US/Western tech industry data — EMEA and APAC teams typically show 15-20% longer cycle times due to timezone-distributed reviews"
  - "AI coding tools inflated individual output metrics (21% more tasks, 98% more PRs merged) but organizational delivery metrics stayed flat — do not confuse individual productivity gains with team-level improvement"
  - "Data from 2025 DORA Report and LinearB 2025 Benchmarks (8.1M+ PRs); market conditions and tooling adoption shift rapidly"

# === SKIP CONDITIONS ===
skip_this_unit_if:
  - condition: "User needs a DevOps maturity assessment, not raw benchmarks"
    use_instead: "business/product-tech/technical-architecture-assessment/2026"
  - condition: "User needs SaaS financial metrics (ARR, churn, CAC), not engineering delivery metrics"
    use_instead: "finance/saas-benchmarks/saas-arr-per-employee-benchmarks/2026"
  - condition: "User needs developer experience or satisfaction benchmarks"
    use_instead: "Search knowledgelib.io for developer experience assessment — no dedicated unit yet"

# === AGENT HINTS ===
inputs_needed:
  - key: segment
    question: "What is the team/org size?"
    type: choice
    options: ["Small team (2-10 engineers)", "Mid-size (11-50 engineers)", "Large (51-200 engineers)", "Enterprise (200+ engineers)"]
  - key: company_stage
    question: "What stage is the company?"
    type: choice
    options: ["Startup (<$5M ARR)", "Growth ($5M-$50M)", "Scale ($50M+)", "Enterprise/Public"]
  - key: metric_focus
    question: "Which metric categories are most relevant?"
    type: choice
    options: ["Velocity (deployment frequency, lead time)", "Stability (change failure rate, MTTR)", "Efficiency (cycle time, throughput)", "Quality (rework rate, PR metrics)", "All metrics"]

# === DISTRIBUTION ===
canonical_source: "https://knowledgelib.io/business/product-tech/engineering-productivity-benchmarks/2026"
suggested_citation: "Source: knowledgelib.io — AI Knowledge Library (verified 2026-03-10, data vintage: Q4 2025)"

# === RELATED UNITS ===
related_kos:
  referenced_by:
    - id: "business/product-tech/technical-architecture-assessment/2026"
      label: "Technical architecture assessment whose Dimension 3 scores CI/CD and deployment maturity on a 1-5 DORA rubric (lead time, change failure rate, MTTR)"
  related_to:
    - id: "business/growth/tech-debt-reduction-playbook/2026"
      label: "Tech debt reduction playbook — quantifying and prioritizing debt against feature delivery, quality gates, sprint allocation"
  depends_on: []
  often_confused_with: []
  alternative_to: []

# === SOURCES ===
sources:
  - id: src1
    title: "DORA State of AI-Assisted Software Development 2025"
    author: Google DORA
    url: https://dora.dev/research/2025/dora-report/
    type: primary_research
    published: 2025-10-15
    data_period: "H1-H2 2025"
    sample_size: "~5,000 technology professionals"
    reliability: authoritative
  - id: src2
    title: "2025 Engineering Benchmarks: Insights from 6.1M+ Pull Requests"
    author: LinearB
    url: https://linearb.io/blog/2025-engineering-benchmarks-insights
    type: primary_research
    published: 2025-09-01
    data_period: "2024-2025"
    sample_size: "8.1M+ PRs from 4,800 engineering teams across 42 countries"
    reliability: authoritative
  - id: src3
    title: "RDEL #115: What are the 2025 benchmarks for the DORA 4 metrics?"
    author: RDEL (Ryan Dahl Engineering Leadership)
    url: https://rdel.substack.com/p/rdel-115-what-are-the-2025-benchmarks
    type: industry_report
    published: 2025-11-01
    data_period: "2025"
    reliability: high
  - id: src4
    title: "Rework Rate is Here: Start Tracking the 5th DORA Metric Today"
    author: Faros AI
    url: https://www.faros.ai/blog/5th-dora-metric-rework-rate-track-it-now
    type: industry_report
    published: 2025-12-01
    data_period: "2025"
    reliability: high
  - id: src5
    title: "Complete Guide to Change Failure Rate [2026 Edition]"
    author: Axify
    url: https://axify.io/blog/change-failure-rate-explained
    type: industry_report
    published: 2026-01-15
    data_period: "2025-2026"
    reliability: high
  - id: src6
    title: "2025 Software Delivery Benchmark Report"
    author: Plandek
    url: https://plandek.com/resources/2025-software-delivery-benchmark-report/
    type: primary_research
    published: 2025-08-01
    data_period: "2024-2025"
    sample_size: "6.1M PRs analyzed"
    reliability: high
---

# Engineering Productivity Benchmarks (DORA + Delivery Metrics)

## Summary

Comprehensive engineering productivity benchmarks covering the five DORA metrics (deployment frequency, lead time for changes, change failure rate, mean time to recovery, rework rate) plus cycle time, PR metrics, and throughput data. Sourced from the 2025 DORA Report (~5,000 respondents) and LinearB's analysis of 8.1M+ pull requests across 4,800 teams. The most significant finding: AI coding assistants boost individual output (21% more tasks completed, 98% more PRs merged) but organizational delivery metrics remain flat — individual productivity does not automatically translate to team-level improvement. [src1]

**Data vintage**: Based on 2025 DORA survey data and LinearB's 2024-2025 PR analysis from 4,800+ engineering teams across 42 countries.
**Key shift**: DORA expanded from 4 to 5 metrics in 2025 by adding rework rate. The framework reorganized into throughput metrics (deployment frequency, lead time, recovery time) and instability metrics (change failure rate, rework rate). The traditional elite/high/medium/low classification was replaced with archetype-based clusters. [src1][src4]

## Constraints
<!-- Agents: read before citing any benchmark number. -->

- These benchmarks represent primarily US/Western tech industry software teams. Do not apply to hardware engineering, manufacturing, or non-software R&D organizations.
- DORA figures are self-reported survey data; measured platform data (LinearB, Plandek) shows materially different distributions. Use DORA for directional comparison, platform data for precision.
- Figures are medians unless otherwise noted. Means are skewed by outlier elite performers; use median for realistic target-setting.
- Data collected H1-H2 2025. If more than 6 months old, search for updated DORA or LinearB reports before citing.
- Only compare teams within the same segment (team size, company stage). A 5-person startup deploying 10x/day is not comparable to a 200-person enterprise deploying daily.

## Metric Category 1: Velocity

### Deployment Frequency

**Definition**: Number of production deployments per unit of time per team/service. Measures how often code reaches production. Counted at the service or application level, not per developer.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| Small team (2-10) | 2-3x/week | 1x/week | 1x/day | Multiple/day |
| Mid-size (11-50) | 1-2x/week | 2x/month | 3-5x/week | 1x/day |
| Large (51-200) | 1x/week | 2x/month | 2-3x/week | Daily |
| Enterprise (200+) | 2-4x/month | 1x/month | 1x/week | 2-3x/week |

**Trend**: Only 16.2% of organizations achieve on-demand deployment (multiple times per day). 23.9% of teams still deploy less than once per month. Distribution is bimodal — teams cluster at high-maturity or low-maturity levels with few in between. [src1][src3]
**Red flag threshold**: Deploying less than once per month indicates batch-oriented delivery with high risk per deployment.
**Action trigger**: If below 25th percentile, investigate deployment pipeline automation, test suite reliability, and change approval bottlenecks.

[src1, src2, src3]

### Lead Time for Changes

**Definition**: Time from code commit to code successfully running in production. Includes code review, CI/CD pipeline execution, and any manual approval gates.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| Small team (2-10) | 1-2 days | 2-5 days | 2-6 hours | < 1 hour |
| Mid-size (11-50) | 2-5 days | 1-2 weeks | 1-2 days | < 1 day |
| Large (51-200) | 3-7 days | 1-4 weeks | 2-3 days | 1-2 days |
| Enterprise (200+) | 1-2 weeks | 1-6 months | 3-7 days | 1-3 days |

**Trend**: Only 9.4% of teams achieve lead times under one hour. 31.9% fall in the one-day-to-one-week range. Anything under 24 hours is a strong result across all segments. [src1][src3]
**Red flag threshold**: Lead time exceeding 1 month signals severe process bottlenecks or manual gates.
**Action trigger**: If lead time exceeds 2 weeks, decompose into sub-phases (coding, review, CI, approval, deploy) to identify the bottleneck — code review is the #1 bottleneck, accounting for 4 of 7 days in average cycle time. [src2]

[src1, src2]

## Metric Category 2: Stability

### Change Failure Rate (CFR)

**Definition**: Percentage of deployments that cause a failure in production requiring remediation (rollback, hotfix, patch, or emergency fix). Does not include planned maintenance or feature flags that are turned off.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| Small team (2-10) | 10% | 15-20% | 5% | < 2% |
| Mid-size (11-50) | 12% | 20-25% | 5-8% | < 3% |
| Large (51-200) | 15% | 25-30% | 8-10% | < 5% |
| Enterprise (200+) | 18% | 30%+ | 10-15% | < 5% |

**Trend**: Only 8.5% of teams achieve the ideal CFR of 0-2%. Elite performers maintain rates below 5%, while high-performing teams stay below 15%. AI-assisted code changes are showing higher initial failure rates due to insufficient testing of generated code. [src1][src5]
**Red flag threshold**: CFR above 25% indicates systemic quality issues — testing gaps, inadequate code review, or poor deployment practices.
**Action trigger**: If CFR exceeds 20%, audit test coverage, review process rigor, and deployment rollback capabilities before increasing deployment frequency.

[src1, src5]

### Mean Time to Recovery (MTTR)

**Definition**: Time from detection of a production failure to full service restoration. Also called "failed deployment recovery time" in the 2025 DORA framework. Measures incident response and remediation capability.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| Small team (2-10) | 1-4 hours | 4-12 hours | 30-60 min | < 15 min |
| Mid-size (11-50) | 2-8 hours | 8-24 hours | 1-2 hours | < 30 min |
| Large (51-200) | 4-12 hours | 12-48 hours | 2-4 hours | < 1 hour |
| Enterprise (200+) | 12-24 hours | 24-72 hours | 4-12 hours | < 2 hours |

**Trend**: Elite teams achieve MTTR under 1 hour across all segments. Recovery time strongly correlates with deployment automation maturity — teams with automated rollback recover 5-10x faster than those requiring manual intervention. [src1][src3]
**Red flag threshold**: MTTR exceeding 24 hours for non-enterprise teams indicates inadequate incident response processes or deployment infrastructure.
**Action trigger**: If MTTR exceeds segment 75th percentile, invest in automated rollback, feature flags, and on-call rotation improvements.

[src1, src3]

### Rework Rate (5th DORA Metric — New in 2025)

**Definition**: Percentage of deployments that are unplanned fixes or patches to correct user-facing defects from prior deployments. Measures delivery instability — how often teams must deploy corrective changes rather than planned features.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| All segments | 8-12% | 15-20% | 4-6% | < 3% |

**Trend**: The 2025 DORA report found that increased AI adoption correlates with increased rework rate — AI-generated code ships faster but requires more post-deployment corrections. Teams with highest AI tool adoption showed measurably higher instability metrics. [src1][src4]
**Red flag threshold**: Rework rate above 15% means the team spends more time fixing production issues than shipping planned work.
**Action trigger**: If rework rate exceeds 12%, audit AI-generated code review practices, test coverage for generated code, and pre-deployment validation steps. [src4]

[src1, src4]

## Metric Category 3: Efficiency

### Cycle Time (PR Open to Merged)

**Definition**: Total elapsed time from pull request creation to merge into the main branch. Includes pickup time (waiting for review), review time, revision cycles, and final approval. Does not include coding time before PR creation or deploy time after merge.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| Small team (2-10) | 3-4 days | 5-7 days | 1-2 days | < 26 hours |
| Mid-size (11-50) | 5-7 days | 7-14 days | 2-4 days | < 2 days |
| Large (51-200) | 7-10 days | 10-21 days | 4-6 days | < 3 days |
| Enterprise (200+) | 10-14 days | 14-30 days | 5-8 days | < 5 days |

**Trend**: Average cycle time across all teams is approximately 7 days, with PRs sitting in the review process for 4 of those 7 days on average. Code review is consistently the single largest bottleneck. Elite teams achieve cycle times under 26 hours. [src2]
**Red flag threshold**: Cycle time exceeding 14 days for non-enterprise teams signals review process breakdown.
**Action trigger**: If cycle time exceeds 7 days, reduce PR size (the #1 driver of cycle time), set review SLAs, and consider automated review assignment. [src2]

[src2, src6]

### Throughput (PRs Merged per Developer per Week)

**Definition**: Number of pull requests merged per developer per week. Measures individual developer output normalized across team sizes.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| All segments | 2-3 PRs/week | 1-2 PRs/week | 4-5 PRs/week | 6+ PRs/week |

**Trend**: Teams using AI coding assistants show 15-25% improvement in PR throughput. However, this increased volume correlates with higher rework rates, suggesting quantity-quality tradeoffs. [src1][src2]
**Red flag threshold**: Sustained throughput below 1 PR/developer/week indicates blockers, context-switching overhead, or oversized PRs.
**Action trigger**: If throughput is low, check PR size distribution first — developers creating large PRs (>500 lines) naturally produce fewer PRs.

[src2]

## Metric Category 4: Quality

### PR Size

**Definition**: Number of code changes (additions + modifications + deletions) per pull request. The single most impactful metric for engineering velocity according to LinearB's analysis.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| All segments | 200-300 lines | 400-661 lines | 100-194 lines | < 100 lines |

**Trend**: Elite teams maintain PR sizes under 194 code changes. Teams with PRs under 194 lines achieve merge frequencies 5x faster than teams with oversized PRs. Larger PRs require more scrutiny, delay approval, and increase error likelihood. [src2]
**Red flag threshold**: Consistently shipping PRs above 500 lines correlates with 3-5x longer cycle times and higher change failure rates.
**Action trigger**: If median PR size exceeds 400 lines, enforce PR size limits, improve task decomposition practices, and adopt trunk-based development. [src2]

[src2]

### Merge Time (Review Complete to Merged)

**Definition**: Time from final code review approval to merge into the main branch. Measures the last-mile efficiency of the delivery pipeline.

| Segment | Median | 25th Percentile | 75th Percentile | Top Decile |
|---------|--------|-----------------|-----------------|------------|
| All segments | 4-8 hours | 12-24 hours | 1-2 hours | < 2 hours |

**Trend**: Elite teams maintain merge times under 2 hours. Long merge times create developer frustration as approved work sits idle. Automated merge queues and CI/CD optimization are the primary drivers of improvement. [src2]
**Red flag threshold**: Merge time exceeding 24 hours after approval indicates CI/CD bottlenecks or manual gate requirements.
**Action trigger**: If merge time exceeds 8 hours, investigate CI pipeline duration, merge queue configuration, and manual approval gates.

[src2]

## Composite Metrics & Rules of Thumb

| Rule | Formula / Threshold | Interpretation |
|------|---------------------|----------------|
| DORA Throughput Score | High deployment frequency + Low lead time | Both must be strong — high frequency with long lead time indicates small, inefficient batches |
| DORA Stability Score | Low CFR + Low MTTR + Low rework rate | All three must be healthy — low CFR with high MTTR means failures are rare but catastrophic |
| Cycle Time Ratio | Review time / Total cycle time < 50% | If review exceeds 50% of cycle time, review process is the bottleneck, not coding or CI |
| PR Size Rule | Median PR < 200 lines | The single highest-leverage metric — drives improvements in cycle time, CFR, and review quality simultaneously |
| Deploy:Rework Ratio | Planned deploys / Rework deploys > 8:1 | Less than 12.5% of deployments should be unplanned fixes; below 5:1 = delivery instability |
| AI Productivity Paradox | Individual output up + Team metrics flat | AI boosts individual velocity but does not automatically improve organizational throughput — requires process adaptation |

**Constraint**: These composite rules apply to software product teams with CI/CD pipelines. Do not apply to data science teams (different workflow), infrastructure teams (different deployment cadence), or teams without automated testing (metrics become unreliable). [src1][src2]

## Segment Definitions

| Segment | Definition | Typical Characteristics |
|---------|-----------|------------------------|
| Small team (2-10 engineers) | Startup or small product team, usually single-service | Direct communication, minimal process overhead, trunk-based development, 1-2 week sprints |
| Mid-size (11-50 engineers) | Growth-stage company or business unit within larger org | Multiple squads, code ownership boundaries emerging, PR reviews required, some specialization |
| Large (51-200 engineers) | Scale-up or division within enterprise | Platform teams, shared services, architecture governance, multiple deployment pipelines |
| Enterprise (200+ engineers) | Large organization or multi-BU software company | Complex CI/CD infrastructure, compliance gates, change advisory boards, cross-team dependencies |

## Year-over-Year Trend Summary

| Metric | 2023 | 2024 | 2025 | Direction |
|--------|------|------|------|-----------|
| Deployment frequency (% daily+) | 30% | 32% | 38% of teams deploy daily or more | Up 8pp over 2 years |
| Lead time (% under 1 day) | 35% | 38% | 41% achieve under 1 day | Up 6pp, steady improvement |
| Change failure rate (median) | 12% | 14% | 15% median across all teams | Up 3pp — AI adoption contributing |
| MTTR (% under 1 hour) | 20% | 22% | 25% recover in under 1 hour | Up 5pp — automation gains |
| Cycle time (average) | 8 days | 7.5 days | 7 days average | Down 12.5% over 2 years |
| PR size (elite threshold) | 250 lines | 220 lines | 194 lines | Down 22% — trend toward smaller PRs |

[src1, src2]

## Common Misinterpretations

- **Treating deployment frequency as the primary metric**: High deployment frequency without stability is counterproductive. A team deploying 10x/day with 30% CFR is worse off than one deploying daily with 3% CFR. Always evaluate throughput and stability together. [src1]
- **Applying enterprise benchmarks to startups (or vice versa)**: A 5-person team should deploy multiple times daily; expecting that from a 300-person org with compliance requirements sets unrealistic targets. Always benchmark within your segment.
- **Equating individual AI productivity gains with team improvement**: The 2025 DORA report found that AI tools boost individual output (21% more tasks, 98% more PRs) but organizational delivery metrics remain flat. The bottleneck shifts from coding to review, testing, and deployment. [src1]
- **Using DORA metrics as targets rather than diagnostics**: Goodhart's Law applies — when deployment frequency becomes a target, teams game it by deploying tiny changes or skipping tests. Use metrics for diagnosis, not incentive compensation.
- **Ignoring PR size while optimizing cycle time**: PR size is the single strongest predictor of cycle time. Teams that focus on review process improvements without addressing oversized PRs see minimal cycle time reduction. [src2]

## When This Matters

Fetch when a user asks about engineering team performance benchmarks, wants to evaluate their DORA metrics against industry peers, is setting engineering KPIs or OKRs, needs to diagnose delivery bottlenecks, or is evaluating the impact of AI coding tools on team productivity.

## Related Units

- [DevOps Maturity Assessment](/business/product-tech/devops-maturity-assessment/2026)
- [Developer Experience Assessment](/business/product-tech/developer-experience-assessment/2026)
- [Technical Debt Assessment](/business/product-tech/technical-debt-assessment/2026)
- [Engineering Team Structure Benchmarks](/business/product-tech/engineering-team-structure-benchmarks/2026)
