Back to blog
12 min read

7 Strategies to Get Real Value from GitHub Analytics as an Engineering Manager

GitHub analytics for engineering managers is only useful when you choose signals that drive decisions and read them with context. This guide shares seven strategies, from deciding what to track to diagnosing flow and risk and interpreting the human side of the numbers, so you can answer whether the team is on track and healthy without turning metrics into surveillance.

7 Strategies to Get Real Value from GitHub Analytics as an Engineering Manager

GitHub holds the most honest record of how your team ships, but that record is easy to misread and easy to drown in. Most engineering managers at startups already have GitHub Insights or a dashboard tool open in a tab, and still can't answer the question their CEO asks every Monday: are we on track, and is the team healthy? The problem is rarely missing data. It is choosing which signals matter, reading them with context, and avoiding the trap of turning metrics into surveillance. These seven strategies for using GitHub analytics as an engineering manager move from deciding what to track, through diagnosing flow and risk, to reading the human side of the numbers.

1. Start with decisions, not metrics

Every number on a dashboard should exist because someone will do something differently when it moves. If no decision hangs on it, it is decoration, and decoration has a cost: it dilutes attention and teaches the team that metrics are theater. Choosing analytics by the decision they support gives each tracked number an owner and an action.

Consider an illustration. A 12-person SaaS team has a 15-chart dashboard that nobody opens after the first week. The engineering manager lists the questions that come up in weekly planning: Is review slowing us down? Are we taking on work we can't finish? Are we shipping changes in chunks too big to review well? She replaces the dashboard with three signals: review wait time, PR age, and merge batch size. Each maps to a question, and each gets discussed on Monday.

To do this in your own team:

  1. List the five management decisions you make repeatedly, such as whom to unblock, what to cut from a sprint, whether to hold a release, and where to add reviewers.
  2. Next to each, name the one GitHub signal that would inform it best.
  3. Drop every metric that does not appear on the list.
  4. Write down who owns each signal and what action a bad reading triggers.
  5. Revisit the list each quarter, because the decisions change as the team grows.

The common mistake is adopting every available metric and then acting on none. Be especially wary of commit counts and lines of code. They are easy to collect and they measure activity, not progress or value. The SPACE framework (Forsgren et al., ACM Queue, 2021) makes the same point: productivity cannot be captured by a single activity metric, and it needs several dimensions read together.

To know it is working, count how many weekly decisions cite a specific signal, and calculate the share of dashboard metrics that triggered an action in the last quarter. Any metric that triggered nothing for a full quarter is a candidate for removal.

2. Break PR cycle time into stages

A single cycle time number tells you that work is slow without telling you where it waits. Cycle time here means the span from a first commit or PR opening to the change being merged or deployed, depending on how you define it. Splitting it into stages shows which wait is actually responsible: time to first review, review rounds, approval to merge, and merge to deploy.

Suppose a team sees a median time to first review of 20 hours, while the coding-to-PR stage takes 6. The instinct might be to push engineers to code faster. The stage data says otherwise: the work sits idle waiting for a reviewer. The right fix is review routing, such as clearer ownership, a review rotation, or a norm of checking the queue at set times, not more pressure on authors.

Putting it in place

  1. Pull PR timestamps (created, review requested, first review, approval, merged) from the GitHub API or an analytics tool. Deployment timestamps may come from your CI/CD system.
  2. Compute the median and the 85th percentile for each stage.
  3. Find the stage with the largest share of total elapsed time.
  4. Set a target for that stage only, and leave the others alone until it improves.

The pitfall: one average

Reporting a single average cycle time hides both outliers and causes. A few PRs that sat for two weeks can drag an average up, while the typical PR is fine; or the average looks healthy while a quarter of PRs wait days. Medians show the typical experience and percentiles show the tail. Averages are rarely the right summary for skewed wait-time data.

Track the median and 85th percentile per stage, week over week. If the median drops but the 85th percentile does not, you have fixed the common case and still have a pocket of work that gets stuck, which is a different problem needing a different fix.

3. Balance review load

Review load is how review work is distributed across the team. In many startups it quietly concentrates on one or two senior engineers, because they know the codebase and everyone requests them by habit. The result is a hidden bottleneck: those seniors have less time for their own work, and every PR in their queue waits.

Imagine a staff engineer who handles 60% of all reviews. Time to first review looks mediocre across the team, but the cause is one overloaded person. The manager adds a CODEOWNERS file that assigns review responsibility by area, sets up a lightweight rotation, and pairs juniors as secondary reviewers on the staff engineer's areas. Time to first review falls, and the juniors build context they would otherwise lack.

To implement it:

  1. Report reviews requested and reviews completed per person over the last 30 to 60 days.
  2. Define CODEOWNERS by directory or service, with at least two owners where possible.
  3. Add a rotation for PRs without a clear owner.
  4. Add a junior as a secondary reviewer, so knowledge spreads without slowing approvals.

The common mistake is rewarding review counts. Once reviews become a score, people approve quickly and shallowly, and quality drops while the numbers look excellent. Use the distribution data to balance work, not to rank people.

Measure the share of reviews handled by your top reviewer, time to first review, and rework after merge, meaning follow-up fixes or reverts of recently merged changes. If the first two improve while rework climbs, reviews are getting faster by getting thinner.

4. Gauge deployment risk with merge volume and churn

Releases fail more often when they carry many changes, large changes, or changes to code that is already unstable. You can estimate that risk before shipping by looking at three things: batch size (merges per release window), churn (how much recently modified code is being modified again), and whether changes touch critical paths. Together these approximate change pressure, the strain a release puts on the system and the team.

Imagine a Thursday release. A pre-release check flags 14 merges touching billing code, and several of the files have been rewritten repeatedly over the past two weeks. The team splits the release, ships the lower-risk changes first, and adds targeted tests around the billing paths. The risk was visible in GitHub data the day before, not after the incident.

To set this up:

  1. Track merges per release window and PR size (lines changed and files touched).
  2. Calculate churn per file over the last 14 days.
  3. Define critical directories, such as payments, auth, and data migrations, and flag any change that touches them.
  4. Compare flagged releases against your change failure rate (the share of deployments that cause a failure needing a fix or rollback, one of the DORA metrics) to calibrate your thresholds.

Be careful with churn. Not all of it is bad: a planned refactor produces high churn on purpose, and early work on a new feature naturally touches the same files repeatedly. Annotate planned work so it does not generate false alarms. The other mistake is waiting for an incident before looking at risk at all.

Progress automates this kind of assessment. It ingests GitHub and Linear data and produces a pre-computed deployment risk read based on merge volume and code churn, so you see the flag without assembling the query yourself. Whether you use a tool or a script, measure the change failure rate, rollbacks or hotfixes per release, and the share of releases flagged high-risk that actually had incidents. That last figure tells you whether your flags are worth trusting.

5. Flag stalled work with age-based signals

Averages smooth over exactly the items that need attention. A PR that has not moved in five days is invisible in a median but obvious in an age-based view. Stall detection works by setting a threshold for inactivity and surfacing every PR or linked issue that crosses it, so someone asks why before the delay compounds.

Suppose a PR has been open for six days, waiting on a reviewer who is on vacation. Without a signal, it surfaces in a retro, if at all. With an automatic flag at day three, the manager reassigns it, and it merges two days later.

A workable setup:

  1. Connect GitHub and your issue tracker, such as Linear, so a stalled PR can be tied to the ticket and initiative it belongs to.
  2. Define stall thresholds by work type. A hotfix might stall at four hours, a normal feature PR at three days, a large refactor at a week.
  3. Post flags daily to Slack or into the standup agenda.
  4. For each flagged item, ask one question: what is it waiting on?

Two failure modes are common. The first is using flags to blame individuals; a stalled item is usually a process or capacity problem, such as unclear ownership, a missing reviewer, or a blocked dependency. The second is setting thresholds so low that the channel fills with noise and the team learns to ignore it. Start conservative and tighten only when the alerts are being acted on.

Progress surfaces stalled work as a pre-computed operational signal across GitHub and Linear, which saves you from maintaining the glue scripts. Measure the count of items stalled beyond threshold and the median time from flag to unblock. The second number shows whether the signal changes behavior or just adds to the pile.

6. Watch momentum trends

Momentum is the direction your delivery is moving: accelerating, steady, or slowing. A single week of data is noisy, but comparing rolling windows of throughput and cycle time reveals trends early enough to act. The aim is to catch a slowdown while it is still a conversation, not a missed milestone.

Imagine a four-week window in which merged PRs fall by 30% while average PR size grows. That pattern often means work is being scoped too large, or that people are holding changes back until they feel complete. The manager raises it with the team, splits the remaining scope into smaller increments, and the milestone stays on track. Nothing dramatic happened in any single week; the trend carried the information.

To read momentum well:

  1. Use rolling four-week medians for throughput (merged PRs or completed issues) and cycle time.
  2. Annotate launches, holidays, incident weeks, and on-call rotations directly on the timeline.
  3. Review the trend monthly with the team, not just privately.
  4. Compare the slope against your plan: is the remaining work achievable at the current pace?

The classic mistake is overreacting to one bad week. A holiday, a big release, or an incident will depress throughput without meaning anything about the team. Annotations and longer windows protect you from false alarms. Progress includes a team momentum analysis that reports whether work is accelerating or slowing, and its Claude-powered Q&A lets you ask in plain language what changed in the last month and get an answer grounded in the underlying activity.

Measure the direction and slope of rolling throughput and cycle time against plan. A steady slope that is still behind plan is a scoping problem, while a falling slope is more often a flow or capacity problem.

7. Pair activity data with morale and workload checks

Delivery metrics describe the work, not the people doing it. Patterns such as sustained late-night or weekend commits, growing friction in reviews, or a workstream that keeps slowing can indicate overload well before it shows up in output or resignations. Treat these as prompts for a conversation, never as verdicts.

Suppose a team shows sustained weekend commits alongside slowing momentum on one workstream. Reading only the slowdown might suggest a performance issue. Read together, the pattern points to overload. The manager brings it up in 1:1s, learns the scope has quietly doubled, and cuts it before anyone burns out.

How to do this responsibly:

  1. Review after-hours activity and review friction (long back-and-forth threads, repeated change requests) at the team level.
  2. Follow up in 1:1s with open questions, not data-backed accusations.
  3. Tell the team exactly what you measure and why.
  4. Keep individual activity data out of performance reviews.

The mistake to avoid is ranking individuals by activity. Leaderboards erode trust and invite gaming, and engineers will optimize for the metric instead of the outcome. Commit counts also say little about who is doing the hardest work. Progress reads team morale and wellness signals alongside momentum, which gives you the human layer next to the delivery data without turning it into a scorecard.

Track the team-level trend in after-hours activity, retention and pulse-survey results, and the number of 1:1 follow-ups actually completed. If the activity signal improves but pulse scores do not, the data is not capturing what people are experiencing, and you should trust the people over the chart.

Where to start, and how the pieces build on each other

Begin with decision mapping and staged cycle time. They cost little, they need only data you already have, and together they tell you what to look at and where work waits. Add review distribution and stalled-work flags next, since both act directly on the bottlenecks you just found and show the team quick, concrete wins. Once those basics have earned the team's trust, layer in deployment risk, momentum trends, and morale checks, which depend on stable definitions and a team that believes the data is used to help, not to judge.

If you would rather have these signals computed and interpreted for you across GitHub and Linear, instead of building and maintaining the queries yourself, that is what Progress is built to do. Learn more about our services


Start your 7-day free trial

Try it on this week's work.

Connect your tools and Progress fills in your last two weeks, so you see what's moving and what's stuck from day one.

7-day free trial · cancel anytime