Hardware fasteners sorted into separate storage bins by type, nuts, bolts and washers

Your Ticket Volume Is a Defect List, Not Demand

Service desks get measured on how fast they close tickets. Almost none get measured on how few arrive.

That sounds like a difference in emphasis. It is not. It quietly decides what the entire function becomes, what gets funded, and which problems in your organisation are allowed to survive indefinitely.

The distinction the framework already makes

ITIL has separated these for decades, and most organisations have read the page.

Incident management restores service. Something is broken for someone right now, and the job is to get them working again as quickly as possible.

Problem management removes the cause. Something has broken repeatedly, and the job is to make sure it stops happening.

They are deliberately separate because they do opposite jobs. Incident management optimises for speed and is measured in minutes. Problem management optimises for permanence and is measured in things that stopped occurring.

In practice, most organisations staff the first one properly and treat the second as something people will get to on a quiet afternoon. The quiet afternoon does not arrive, because a queue that is never emptied always has something in it.

So the queue becomes permanent. And once it is permanent, it gets managed rather than emptied. Rotas, SLAs, dashboards, headcount models. All of it sophisticated, all of it about handling volume rather than removing it.

The diagnostic: read the queue by category, not by ticket

Here is the exercise, and it takes about twenty minutes.

Export a month of first-line tickets. Ignore the individual tickets entirely. Group them by category, then sort by count. Now look at the top five.

In most organisations they are not a surprise, and they are not varied. They are the same handful of things every single month:

  • An access request that should have been granted automatically at joiner stage
  • A piece of software that fails the same way on the same build
  • A certificate, a licence or a domain with no named owner for its renewal
  • A password or MFA reset caused by a process nobody has revisited in three years
  • A manual step that exists only because a human has to translate one system into another

That is not demand. That is a defect list. And it is being worked as though it were demand, by people who are rated on how quickly they work it.

The distinction matters because the two require completely different responses. Demand needs capacity. Defects need engineering. Treating the second as the first is how a service desk grows without anything getting better.

The uncomfortable part

A service desk that gets very good at closing tickets removes the pressure that would have fixed the cause.

Speed becomes an anaesthetic. If the top category is resolved in four minutes by a competent first-line team, nobody upstream ever feels it. The cost is real, it is just distributed into a place where it looks like a staffing line rather than an engineering problem. The better the team performs, the more invisible the underlying fault becomes.

This is why “our service desk is performing well” and “our environment is full of unfixed defects” are not contradictory statements. They are frequently the same statement.

And when it is not fast enough, something worse happens

Once the desk is understood to be a bottleneck, people stop using it.

Not as a protest, and not all at once. They just start solving things themselves, because raising a ticket and waiting feels worse than having a go. Someone finds a free tool that does the job. Someone shares the file the quick way. Someone keeps a spreadsheet that has quietly become a system of record for a process the business now depends on.

That is most of what shadow IT actually is. Rarely defiance. Usually people routing around a queue they have learned not to trust, with wildly varying degrees of success and no way for you to tell which.

And it comes back to you anyway. It arrives the day it breaks, or the day the person who built it leaves, and by then you are supporting something you did not choose, did not design, cannot see and have no documentation for. The workaround has become tech debt, and the tech debt will generate its own tickets.

That is the loop:

  1. A defect is absorbed by the service desk rather than fixed
  2. The queue stays full, so the desk is slow
  3. People route around the slow desk and build their own solutions
  4. Those solutions land back on IT unsupported, as new tech debt
  5. The new tech debt generates its own tickets

Every turn of that loop makes the next turn worse, and every stage of it looks locally rational to the person doing it.

The same loop runs through access

This one is expensive enough to deserve its own section, because it converts a service delivery problem into a security problem.

If getting the right permission takes three days, people ask for everything they might conceivably need while they have someone’s attention. Just-in-case access instead of just-in-time. Zero trust quietly becomes “I might need it”.

And none of it comes back, because nobody has ever raised a ticket to have something taken away from them.

Note what that does to your numbers. Access sprawl grows while the ticket count falls, because people are asking less often by asking for more. On the only metric anyone is watching, that looks like an improvement.

If that pattern is familiar, it is the same one that surfaces the moment anyone tries to deploy an AI agent against a tenant. The permissions were granted for convenience years ago and nobody revisited them. We covered that in detail in the agent access and permission sprawl pillar, with the joiner-mover-leaver mechanics in the JML piece and the discipline itself in least privilege in practice.

It also poisons the data you use to argue for investment

Here is the part that makes this hard to escape once you are in it.

Problems people solve themselves never get logged. They never appear in the volume you present when you ask for headcount, tooling or budget. So the queue is not just a defect list, it is a partial defect list, filtered down to the problems people still believe are worth reporting.

Which means the areas where your service is worst are systematically under-represented in your own evidence. The worse a service gets, the fewer tickets it generates, because people give up on it. Any measure built purely on ticket volume will therefore point you at the areas people still have faith in.

What this article is not claiming

You cannot engineer a queue to zero, and anyone selling you that is selling you something.

New starters need accounts. New kit needs setting up. People meeting something for the first time need help, and helping them is the job rather than a failure of design. External failures happen. Some things are genuinely cheaper to handle with a human than to automate, and that remains true no matter how good your tooling gets.

The claim is narrower: some of your volume is demand and some of it is the same defect arriving repeatedly, and most organisations never separate the two. Once you sort by category rather than by date, the two are not difficult to tell apart.

Measure something else alongside

If you take one practical thing from this, it is that throughput metrics are not wrong, they are just insufficient on their own. They describe how well you absorb. Nothing in them describes whether you should have had to.

Worth tracking alongside the usual set:

  • Share of volume from your top five categories. If five categories are most of your queue, most of your queue is fixable rather than inevitable.
  • Repeat rate per category, month on month. A category that is stable for six months is not demand, it is a decision nobody has made.
  • Time from a category being identified to the cause being addressed. Most organisations cannot produce this number at all, which is itself the finding.
  • Categories that fell to zero. The only metric on this list that measures problem management doing its job, and the only one that ever gets celebrated less than it should.

The question worth asking

Every organisation has someone who owns the queue. Far fewer have anyone who owns the reasons.

So the question is not whether your service desk is efficient. It is whether anybody is reading that queue as evidence, or simply working it.

Image: Joaquin Reyes Ramos via Pexels.

Enjoyed this guide?

New articles on Linux, homelab, cloud, and automation every 2 days. No spam, unsubscribe anytime.

Scroll to Top