Maximize IT Value With Top Performer Operational Practices
There's a moment in nearly every IT leader's year when the budget conversation turns uncomfortable. You're sitting across from a CFO or a board committee, and the question lands: "We've spent $4 million on technology this year. What did we get?"
If your answer is a list of projects delivered and tickets closed, you're probably going to lose that conversation. Not because your team did bad work, but because the numbers you're quoting describe activity, not value. And activity is easy to dismiss.
Here's the thing that took me a long time to internalize: the gap between an IT organization that's constantly firefighting and one that's quietly reliable isn't mostly about talent or budget. It's about operational practices. Specifically, the small set of disciplines that top performer operational practices tend to share, regardless of industry, company size, or which cloud vendor they happen to use.
That's the thesis of this article. We're going to walk through what those practices actually look like in practice, why so many organizations skip them, how to tell whether you're making real progress, and where to start if you're staring at a mess and don't know which end to grab. Along the way, I'll point to research from the IT Process Institute (ITPI), which has been studying high-performing IT organizations since 2004 and has turned that research into a fairly practical body of work.
Grab a coffee. This one's long, but it's the kind of long that saves you a year of wandering.
---
Why Good IT Teams Still Struggle to Show Value
Let me start with a pattern I've seen repeatedly.
A mid-sized healthcare organization has a competent IT department. Maybe 60 people. They've adopted a ticketing system, they do change reviews, they run a decent monitoring stack. Nothing is obviously broken. And yet:
- Roughly a third of their engineering time goes to unplanned work — outages, emergency fixes, things that broke that shouldn't have.
- Their change success rate hovers around 85%, which sounds fine until you realize their peer group is running at 98%+.
- Nobody can answer, with confidence, how many production servers they have, what version of the database is running where, or which systems still depend on a vendor that was acquired three years ago.
- They just bought their fourth security tool, and their incident count hasn't moved.
This isn't a story about incompetence. It's a story about a missing operating model.
Most IT organizations grow by accretion. You add a tool here, a process there, a new team when the workload justifies it. Ten years later you have a system that works, sort of, but has no coherent design. It's the equivalent of a house where every addition was built by a different contractor with a different idea of where the load-bearing walls are.
The problem compounds because complexity itself is expensive. Every additional configuration item, every undocumented dependency, every "tribal knowledge" workaround increases the chance that the next change breaks something unrelated. And when that happens, the fix eats time that was supposed to go toward the strategic project the business actually asked for.
So you end up in a loop: too much unplanned work → less time for improvement → more fragile systems → more unplanned work.
Breaking that loop is what top performer operational practices are designed to do. Not through heroics, and definitely not through buying another platform. Through discipline applied in a specific order.
---
What "Top Performer" Actually Means in IT Operations
The phrase gets thrown around loosely, so let's define it. A top performer isn't the company with the biggest IT budget or the most fashionable tech stack. It's the organization that, relative to peers of similar size and complexity, delivers:
- Higher change success rates. Changes go in, work as intended, and don't cause incidents.
- Lower unplanned work as a share of total effort. Their calendars aren't dominated by reacting.
- Faster mean time to restore (MTTR) when something does break.
- Better first-contact resolution on the service desk.
- Predictable delivery — release dates they actually hit, without death marches.
Notice that none of those are technology metrics. They're all operational outcomes. Which is the whole point.
How ITPI studies top performers
The IT Process Institute was founded in 2004 as an independent research organization specifically to figure out which practices actually differentiate high performers from everyone else. The methodology is comparative: instead of surveying the industry and reporting the average, ITPI looks at the top decile of performers and asks what they do differently, then translates the answers into prescriptive guidance.
That distinction matters more than it sounds. Most analyst research tells you what the market is doing. ITPI's work tells you what the best are doing and how to copy them. That's a different product entirely, and it's the reason their material reads less like a market report and more like an instruction manual.
The two numbers that reveal almost everything
If you only had access to two metrics from an IT organization, you could predict a lot about its health.
1. Change success rate. What percentage of changes to production are completed without causing an incident or requiring a rollback?
2. Unplanned work ratio. What share of your team's total effort goes toward unplanned, reactive work?
The second number is the one that surprises people. In the original research behind the Visible Ops work, ITPI found that a striking majority of unplanned outages — roughly 80% — traced back to changes the organization itself had made. Not hardware failures. Not acts of God. Changes. That finding reframed the entire conversation: if most of your downtime is self-inflicted, then most of your downtime is preventable through better change discipline.
Sit with that for a second. Eighty percent of your bad days are things you did to yourself.
---
The Core Practices That Separate Top Performers From the Rest
Once you know where the pain comes from, the fix becomes less mysterious. Top performers tend to invest in four practice areas, and they get them working in roughly this order.
Change management that people actually follow
Most organizations have a change management process. Few have one people respect.
The difference shows up in the details. Top performers:
- Classify changes by risk, not by volume. A standard, pre-approved change (restarting a service, applying a documented patch) moves fast with minimal ceremony. A high-risk change gets more eyes.
- Keep the approval path short. If getting a change approved takes three days, engineers will route around the process. Guaranteed.
- Track change success rate as a first-class metric and review it monthly.
- Treat every failed change as a learning event, not a disciplinary one.
That last point is where most organizations quietly fall apart. If people get punished for surfacing problems, they stop surfacing them. Then your incident data becomes fiction.
Configuration management and knowing what you own
You cannot control what you can't see. Configuration management is the practice of maintaining an accurate, current record of your infrastructure, applications, and the relationships between them.
This is the least glamorous work in IT, and it's the most valuable. Consider what becomes possible once you actually know your environment:
- Impact analysis before a change, so you know what you'll break.
- Faster root cause analysis during incidents, because you can trace dependencies.
- Accurate disaster recovery planning, since you know what needs to come back and in what order.
- Real security posture, because you can find the unpatched system everyone forgot about.
The standard advice is to start with the highest-value, highest-risk systems, not with everything. Trying to document your entire estate in one push is how configuration management projects die. Pick twenty systems that matter most, get them accurate, and expand from there.
Release management as the middle layer
Between "I changed a config" and "we shipped a new version of the application" there's a whole discipline most organizations underinvest in: release management. This is the practice of building and deploying software in a repeatable, predictable way.
Top performers build what's sometimes called a repeatable build library — standardized, tested, documented ways to construct and deploy each of their applications. When you have that, rebuilding an environment stops being an archaeology project and becomes a checklist.
The payoff shows up in two places. First, deployments get boring, which is exactly what you want. Second, when disaster recovery is actually needed, you can rebuild from the library instead of from memory.
Monitoring, incident response, and post-incident reviews
Everyone monitors. Top performers monitor with intent and act on what they find in a structured way. That means:
- Alert thresholds tuned so that alerts mean something (alert fatigue is a real productivity killer).
- Defined severity levels with matching response expectations.
- Blameless post-incident reviews that produce specific, assigned follow-up actions.
- A feedback loop where recurring incidents trigger a permanent fix rather than the same band-aid applied for the ninth time.
One quick test: pull your last ten incidents. How many of them share a root cause with an earlier incident? If the answer is more than two, you don't have an incident response problem. You have a follow-through problem.
---
Visible Ops: A Framework Built From Evidence
All of the above gets organized, in ITPI's work, under a framework called Visible Ops. If you've been in IT for a while, you've probably seen the orange book on someone's shelf. The Visible Ops Handbook has sold over 400,000 copies, which for a book about IT operations is an absurd number. It's popular because it's short, specific, and written for people who have to do the work on Monday, not write a dissertation about it.
The four steps
The Visible Ops approach walks through four sequential phases. The order matters — skipping ahead is a common reason improvement efforts stall.
Step 1: Stabilize the patient. Before you build anything new, get control of change. Implement a lightweight change management process, establish a change advisory function that doesn't become a bottleneck, and start tracking change success rate.
Step 2: Find and fix fragile artifacts. Identify the systems that cause disproportionate pain — the ones that break often, that nobody understands, that have undocumented configurations — and fix them. Often this means documenting them, standardizing them, or retiring them.
Step 3: Create a repeatable build library. Once the environment is stable and visible, codify how things get built so that builds become repeatable rather than artisanal.
Step 4: Continuously improve. With stability and repeatability in place, use measurement to keep tuning.
Notice that step one isn't "buy a tool." It's "stop the bleeding." That sequencing is deliberate and, in my experience, correct.
What the book series covers
ITPI extended the Visible Ops methodology over the years to cover areas where the same discipline applies but the specifics differ. The series now includes:
- The Visible Ops Handbook — the core IT operations playbook.
- Visible Ops Security — applying the methodology to security operations.
- Visible Ops Private Cloud — how high performers build and operate private cloud environments.
- Visible Ops Cybersecurity — governance, culture, and controls for security programs.
- VisibleOps A.I. — the newest addition, covering AI governance and operations. It debuted at #25 on Amazon's Top 100 New Releases in Computers & Technology, which suggests the timing was right.
You don't have to read them in order, but if you're new to the framework, the original handbook first is the sensible move. Everything else builds on it.
---
Cloud and Private Cloud: Value Beyond Migration
Let's talk about cloud, because this is where a lot of organizations have spent serious money with mixed results.
The pattern I see most: a company decides to move to cloud, sets a "migrate X% of workloads by [date]" goal, and hits it. Then eighteen months later, the bill is 40% higher than projected, performance is inconsistent, and nobody's entirely sure which workloads should be where.
Migration isn't the goal. Value is the goal, and migration is only one way to get there.
Why lift-and-shift underdelivers
Lifting an application from your data center to a cloud provider without changing how you operate it usually produces one of two outcomes:
- You pay more for the same thing. You've moved the workload but kept the operational practices, so you've added cloud cost without cloud benefit.
- You pay less but lose control. Auto-scaling kicks in unexpectedly, costs spike during peak, and your change management process doesn't account for infrastructure that changes itself.
Neither is a disaster, but neither is the outcome anyone promised in the business case.
What top performers do differently
ITPI's private cloud research looked at organizations that were actually getting value from cloud investments, and the differences were operational, not technical:
- They define a service catalog. Engineers request from a menu of approved, pre-configured services instead of building bespoke environments.
- They standardize the build. A new environment gets built from a template, not from a wiki page last updated in 2019.
- They measure unit economics. Cost per transaction, cost per customer, cost per environment — whatever fits the business — and they review it.
- They treat cloud resources as configuration items subject to the same change and configuration discipline as everything else.
- They keep governance light but real. Guardrails, not gates.
If you're mid-migration right now, the highest-leverage thing you can probably do this quarter is define a service catalog with a handful of approved patterns. It sounds basic. It changes everything downstream.
---
Cybersecurity: Governance, Culture, and Controls in the Right Order
Security is where the stakes get highest and the discipline gets weakest.
The typical failure mode goes like this: a breach happens somewhere in the industry, the board asks questions, leadership responds by buying tools. New EDR platform, new SIEM, new DLP, new something. Six months later, nobody's sure whether the organization is any safer.
Tools matter. They're just not the first thing.
The order that works
Based on ITPI's security research, high performers tend to build in this sequence:
- Governance. Who owns security outcomes? What's the risk appetite? What gets reported to the board, and how often?
- Culture. Do people report suspicious activity without fear? Do developers flag security concerns early, or do they get overridden by delivery pressure?
- Process. How are incidents detected, triaged, escalated, and reviewed? Who has authority to act?
- Technical controls. Now the tools, deployed against a real understanding of what you're protecting.
Flip that order — which is what most organizations do — and you get expensive tooling sitting on top of unclear ownership and a culture that treats security as somebody else's job.
Compliance is a side effect, not the goal
This one's worth stating plainly because it trips up a lot of people. If you build good security practice — accurate asset inventory, controlled changes, tested incident response, reviewed access — you'll pass most audits without much extra effort. The documentation auditors want is a byproduct of doing the work.
If instead you optimize for passing the audit, you'll build a paper trail that satisfies the checklist and does very little to reduce actual risk. Then you'll fail the next audit anyway, because the practices weren't real.
Visible Ops Cybersecurity digs into this, particularly around the governance, training, and culture side that technical controls can't compensate for. There's a reason ITPI frames it as a holistic problem. Security is one of the few areas where the org chart matters as much as the architecture diagram.
---
DevOps and Automation Without Breaking Production
DevOps gets treated as a tooling category more than a practice. Buy a CI/CD pipeline, adopt containers, hire a platform team, done. Except that's not done at all.
DevOps, at its core, is about integrating the people who build software with the people who run it. The tooling supports that integration. It doesn't create it.
Automate the right things in the right order
Automation is a force multiplier on whatever process you already have. If the process is good, automation makes it fast. If the process is broken, automation makes it break faster and at scale.
So the sequence matters:
- Standardize the process before automating it. If three teams deploy three different ways, you don't have a pipeline problem. You have a standards problem.
- Automate the build first. Repeatable builds are the foundation everything else rests on.
- Then automate testing. Especially the tests that catch the failures you actually experience.
- Then automate deployment. And connect deployments to your change management process rather than exempting them from it.
- Then automate the feedback loop. Monitoring, alerting, and rollback.
A deployment pipeline that bypasses change control isn't progress. It just means changes are happening invisibly, which makes the next incident harder to diagnose.
The measurement question
The DevOps world has good public research on delivery performance — deployment frequency, lead time for changes, change failure rate, time to restore. Those four metrics are useful, and they overlap neatly with the operational metrics discussed earlier. Change failure rate and change success rate are two sides of the same coin.
What I'd caution against is treating the metrics as targets. Deployment frequency isn't inherently good. Deploying twenty times a day with a 30% failure rate is worse than deploying twice a month with a 2% failure rate. Unless you're a company where daily releases are the actual business requirement, the goal is reliable delivery at whatever cadence the business needs.
---
AI Governance: Applying Proven Practice to a New Problem
Artificial intelligence is the newest pressure on IT organizations, and the response has been a mix of enthusiasm, paralysis, and shadow adoption.
Here's the pattern: employees start using AI tools before any policy exists. Some of those tools handle sensitive data. Leadership eventually notices, forms a committee, and then spends six months debating a governance framework while adoption continues underneath.
The good news is that the fundamentals of IT governance don't change just because the technology is new. What changes is the specifics.
The practices that carry over
- Know what you have. Inventory AI systems in use, including the ones individuals signed up for on their own. You can't govern a system you don't know exists.
- Control changes. When a model gets updated, a prompt gets modified, or data sources change, that's a change. Treat it like one.
- Classify data first. The most important governance question is what data is going into the model. Everything else follows from that.
- Define ownership. Who's accountable when an AI system produces a bad output? If nobody can answer that, you don't have governance.
- Measure outcomes, not adoption. "We deployed 12 AI use cases" is not a value metric. Time saved, error reduction, or revenue impact is.
Where VisibleOps A.I. fits
ITPI's newest publication takes the same research-driven approach and applies it to AI operations. It's aimed at the leaders who need to move from "we should probably do something about AI" to an actual operating model — without reinventing a governance framework from scratch when the principles from change management and configuration management already apply.
If you're in that spot, the book is a faster path than building your own framework through trial and error. You can find it alongside the rest of the series through the ITPI store.
---
Measuring What Matters: Metrics That Reflect Real Value
Metrics are where good intentions go to die. Organizations either measure nothing, or measure forty things, and in both cases the data doesn't drive decisions.
Here's a workable set. Pick five to start.
| Metric | What it tells you | Rough target for a solid performer |
|---|---|---|
| Change success rate | Whether changes are safe | 95%+ |
| Unplanned work ratio | How much time goes to reacting | Under 25% |
| Mean time to restore (MTTR) | How fast you recover | Trending down quarter over quarter |
| Repeat incident rate | Whether you actually fix root causes | Under 10% |
| Configuration accuracy | Whether your inventory is trustworthy | 95%+ on in-scope systems |
| Percentage of changes that are standard/pre-approved | Process maturity | 50%+ |
| Patch compliance within SLA | Security hygiene | 95%+ |
A few notes on using these.
Measure trends, not snapshots. A single month's number is noise. Three months of movement is signal.
Don't measure anything you won't act on. Every metric you track consumes attention. If a metric has never triggered a decision, stop reporting it.
Watch for gaming. If change success rate becomes a performance target for individuals, expect people to reclassify failed changes as something else. Metrics work at the team and system level, not the individual blame level.
Pair leading and lagging indicators. MTTR is lagging — it tells you about problems that already happened. Percentage of standard changes is leading; it predicts future stability.
---
Culture and Leadership: The Part Nobody Can Buy
I've saved this for near the end on purpose, because it's the part that's hardest to implement and most likely to determine whether anything else works.
Process changes fail when the culture doesn't support them. Specifically:
- If leadership treats IT as a cost center to be minimized, the team will optimize for looking cheap rather than being effective.
- If post-incident reviews turn into blame sessions, incident reports become sanitized and useless.
- If managers are rewarded solely for project delivery, they'll take shortcuts on operational hygiene to hit dates.
- If frontline engineers aren't involved in designing the process, the process will be designed for how leaders imagine work happens rather than how it actually happens.
None of these are exotic insights. They're just rarely acted on.
What top performers do about culture
ITPI's research consistently finds that high-performing IT organizations treat culture, leadership, and process as interdependent rather than separate concerns. Practically, that shows up as:
- Executive sponsorship that's visible. Someone at the top talks about operational discipline in the same breath as delivery.
- Blameless reviews with real follow-through. Not "blameless" as a slogan, but as an enforced norm where the questions are about systems, not people.
- Engineers in the room when processes get designed. The people doing the work know where the friction is.
- Middle managers protected from contradictory incentives. You can't demand both 100% project delivery and rigorous change discipline without giving managers room to say no.
- A tolerance for the slow start. The first 90 days of a change management program often look like reduced velocity. That's normal. The payoff comes later.
A note on the "we're too small for this" objection
I hear this a lot from smaller IT teams. It's understandable, and it's wrong.
A 15-person IT department needs change discipline more than a 500-person one, because a single person is a larger share of the capacity. When one outage eats 20% of your team for two days, that's a much bigger proportional hit. The practices scale down fine. You just apply them with less ceremony.
---
Common Mistakes That Undo IT Improvement Efforts
Let's go through the failure patterns. If you recognize your organization in three or more of these, you're not alone — and knowing which one is biting you is half the battle.
1. Starting with tools. Buying a platform before defining the process. The tool will enforce whatever you configure it to enforce, including your existing dysfunction.
2. Boiling the ocean. Trying to document every server, classify every change, and instrument every service in one push. This always stalls around month three. Start narrow.
3. Making change management a paperwork exercise. If the process consists of filling out forms nobody reads, engineers will route around it, and you'll have lost the data you need.
4. Skipping configuration management because it's boring. Also the most common reason improvement efforts plateau. You can't do impact analysis on an environment you can't see.
5. Copying a top performer's process without context. A practice that works at a company with 3,000 engineers may be absurd at a company with 40. Borrow principles, adapt specifics.
6. Treating compliance as the goal. Discussed above, but it bears repeating because it's so seductive — passing the audit feels like progress.
7. Measuring everything. Forty dashboards, no decisions. Fewer metrics, more action.
8. No executive owner. Operational improvement without sponsorship becomes a side project for whoever cares most. Then that person changes jobs, and it dies.
9. Quitting at the 90-day mark. The first quarter of any operational discipline program feels worse before it feels better. Teams that push through months four through nine see the payoff; teams that quit at day 90 restart from zero in two years.
10. Ignoring the security and AI dimensions. If your operational improvement plan doesn't extend to how you govern AI systems and security controls, you'll end up building a second, disconnected process later.
---
A 90-Day Roadmap to Move Toward Top Performer Practices
Enough theory. Here's a concrete sequence if you're starting from scratch or restarting a stalled effort.
Days 1–30: Get visibility and establish baseline
- Pick five to ten production metrics and start measuring them. Change success rate and unplanned work ratio first.
- Pull your last 20 incidents. Categorize them by root cause. Identify repeats.
- Inventory your top 20 business-critical systems. Just twenty — resist the urge to do more.
- Identify who currently approves changes and how long approval takes. Measure it.
- Get executive sponsorship on the record. One named owner at the leadership level.
Days 31–60: Stabilize change
- Implement a two-tier change process: standard (pre-approved, low risk) and normal (reviewed, scheduled).
- Define what qualifies as standard. Start with ten change types you do all the time.
- Set a target approval time for standard changes — under an hour, ideally.
- Start tracking change success rate weekly and share the number with the whole team. No blame, just visibility.
- Begin configuration documentation on those top 20 systems. Assign an owner per system.
Days 61–90: Build the repeatable layer
- Identify the top three systems you rebuild or redeploy most often. Document the build.
- Formalize post-incident reviews. Blameless, with assigned follow-up actions and due dates.
- Stand up a lightweight dashboard with your five core metrics.
- Run a retrospective on the 90 days: what's working, what's friction, what's being ignored.
- Set the next quarter's target — usually, get change success rate above 90% and configuration coverage to 50% of critical systems.
Months 4–12: Extend and deepen
- Expand configuration management to the next tier of systems.
- Formalize release management with a build library for your main applications.
- Extend change discipline to cloud resources and AI systems.
- Run a tabletop incident response exercise.
- Consider the Visible Ops series for the areas you're weakest in — security, private cloud, or AI.
If you want a head start on any of these phases, ITPI's research publications and executive snapshots are designed for exactly this use — practical enough to hand to a working team, not just a leadership deck. They're not the only resource out there, but they're unusually specific, and the shared research funding model keeps them affordable compared to traditional analyst subscriptions.
---
Frequently Asked Questions
How long before we see measurable results from operational improvements?
It depends on where you're starting, but the typical pattern is: change success rate improves within 60 to 90 days (it's the fastest lever), unplanned work starts declining around month four to six, and MTTR trends down over six to twelve months. Culture shifts take longer — often a year or more — because they depend on repeated evidence that the new way works.
Do we need a big budget or a specific toolset to do this?
No. Almost nothing described in this article requires new spending beyond maybe a documentation tool and some time. The Visible Ops methodology was explicitly designed to be implemented with existing staff and existing tools. In fact, buying tools first is one of the most common ways these efforts fail. If you have budget, spend it on training and research before platforms.
How is ITPI's research different from a big analyst firm?
Three differences stand out. First, ITPI focuses specifically on comparing top performers to everyone else rather than reporting market averages. Second, the output is prescriptive — step-by-step guidance rather than trend analysis. Third, the shared research model means costs are spread across participating organizations, which makes it considerably cheaper than traditional analyst subscriptions. If you need a broad market overview, an analyst firm is the right call. If you need to know what the best organizations do and how to copy them, ITPI is more direct.
We're a small IT team. Is this realistic for us?
Yes, and it's arguably more important. The practices scale down. A five-person team might have ten standard change types instead of a hundred, and a configuration inventory of fifteen systems instead of fifteen thousand. The core discipline — know what you have, control changes, learn from incidents — applies at every size.
What's the relationship between Visible Ops and ITIL?
They're complementary. ITIL provides a broad service management framework covering a wide range of processes. Visible Ops is narrower and more prescriptive, focused on the specific operational practices that differentiate top performers. Many organizations use ITIL for vocabulary and structure, then use Visible Ops for the operational specifics. You don't have to pick one.
How do we start governing AI without slowing down adoption?
Start with three things: an inventory of AI tools already in use, a data classification policy that says what data can and can't go into external models, and a named owner for AI governance. That's enough to prevent the worst outcomes while you build out a fuller framework. VisibleOps A.I. walks through the rest in more detail if you want a structured path.
Should we measure individual performance against these metrics?
Generally, no. Change success rate, unplanned work ratio, and similar metrics work at the team and system level. When you attach them to individual performance reviews, people start gaming the definitions — reclassifying failed changes, mislabeling unplanned work as planned. The purpose of measurement is learning, not ranking.
What if we've already tried this and it stalled?
Find out where it stalled. Usually it's one of three places: the process became bureaucratic and people routed around it, the configuration management effort lost steam because it was scoped too broadly, or the executive sponsor changed roles and the mandate disappeared. Fix that specific failure rather than restarting the whole program. Restarting from zero is the most expensive option available.
---
Where to Go From Here
Let's bring this back to the opening question: "We spent $4 million on technology. What did we get?"
The answer, for a top-performing IT organization, is a set of operational outcomes — high change success, low unplanned work, fast recovery, predictable delivery, defensible security, and governed AI adoption. Those outcomes don't come from any single purchase. They come from a small number of practices applied consistently, in the right order, with leadership that treats operational discipline as a real priority rather than a slogan.
If you take one thing from this article, take the sequencing. Stabilize change first. Then make your environment visible. Then build repeatability. Then optimize. Skipping ahead to step four is the most common and most expensive mistake.
And if you'd like a structured path rather than building the framework yourself, that's essentially what the IT Process Institute exists to provide. Their research studies and benchmarking reports show you what top performers do; the Visible Ops book series — including the newer titles on security, private cloud, cybersecurity, and AI — translates that research into steps a working team can follow. You can browse the full catalog, along with executive snapshots, eBooks, and training webinars, through the ITPI store.
Start with the two metrics. Track them for a month. Then pick one change — one — and fix it. That's the whole trick, repeated for a few years.
One more thing: if you're reading this and thinking "our situation is messier than what's described here," that's fine. It almost always is. The mess isn't the problem. The absence of a sequence for cleaning it up is. Now you have one.
