Reliability Tab
The Reliability tab surfaces failure patterns, recovery speed, and non-deterministic behaviour in your CI workflows.
KPI Cards
| Metric | Definition | Target |
|---|---|---|
| MTTR | Mean Time To Recovery — the average time between a failing run and the next successful run on the same branch. Measures how quickly the team resolves CI breakages. | < 1 hour (DORA Elite) |
| Failure Streak | Number of consecutive failed runs on the default branch with no successful run in between. Any streak ≥ 3 is flagged red. | 0 — any streak warrants immediate attention. |
| Flaky Branches | Number of branches where the last 10 runs alternated between success and failure (flip-flop pattern). Indicates non-deterministic tests or environment instability. | 0 flaky branches |
| Re-run Rate | Percentage of runs that were manually re-triggered (run_attempt > 1). A high rate is a strong signal of flaky tests or infrastructure instability. | < 5% |
Charts
| Chart | What it shows | How to read it |
|---|---|---|
| Pass / Fail Timeline | Bar chart of run outcomes ordered chronologically. Green bars = success (+1), red bars = failure (−1). | Clusters of red reveal outage duration and frequency. Hover a bar to see the run number and conclusion. |
| Flaky Branches | Badge list of branches that exhibited flip-flop outcomes in the last 10 runs. | Any branch listed here has non-deterministic CI. Investigate and quarantine unstable tests. |
| Anomaly Detection | Runs whose duration deviated more than 2 standard deviations from a rolling 10-run baseline. | Classified as moderate (2–3 stddev) or extreme (> 3 stddev). Investigate for stuck jobs, infrastructure issues, or abnormally large changesets. |
CI / Workflow Metrics
| Metric | Definition | Good range |
|---|---|---|
| Success Rate | Percentage of completed workflow runs that finished with conclusion = success over the selected window. | > 90% |
| Duration P95 | The 95th-percentile run duration. 95% of runs complete faster than this value. A rising P95 indicates flaky or slow test suites. | Depends on workflow type; watch for upward trend. |
| Queue Wait P95 | The 95th-percentile time between a run being triggered and its first job actually starting. High values indicate runner capacity constraints. | < 2 minutes for self-hosted; < 5 min for GitHub-hosted. |
| Avg Queue Wait | Mean runner wait time across all runs. This is pure infrastructure overhead — not code or test execution time. | < 1 minute ideally. |