Blog

Five things we find in every Intune health check

Every tenant is different; the findings rarely are. These are the five problems that surface in almost every Intune health check we run, why they happen, and where to start fixing each one.

Most Intune tenants were configured during a project: a Windows 11 migration, an Autopilot rollout, a rushed response to a security review. The project ends, the team moves on, and the tenant is left running on decisions nobody remembers making. When we come in to health-check an estate, the same five findings appear again and again, regardless of sector or size.

1. Update rings that exclude half the estate

The rings exist, they are sensibly staged, and the pilot ring even has the right people in it. But the assignments were scoped during the pilot and never broadened, or a dynamic group query stopped matching after a naming convention changed, or an exclusion group quietly grew every time an update caused a problem. The result is a large population of devices that belong to no ring at all, patching themselves on default behaviour or not at all.

The risk is subtle because reporting looks healthy: the devices in rings report green, and the devices outside them simply do not appear. With Windows 10 support now ended and monthly quality updates carrying the security load, an invisible unpatched population is exactly the exposure a board assumes it does not have.

First fix: reconcile ring membership against the full device inventory. Every managed device should resolve to exactly one ring. Build a report of the unassigned remainder and drive it to zero before touching anything else.

2. Compliance policies that evaluate nothing

Intune ships with a tenant-wide default that decides what happens to a device with no compliance policy assigned. Left on its out-of-the-box setting, those devices are marked compliant. Combine that with policies assigned to user groups that miss shared and kiosk devices, or a policy whose only rule is that a policy exists, and you get a dashboard full of compliant devices where nothing was ever actually checked.

A dashboard full of compliant devices where nothing was ever actually checked.

This matters because compliance is rarely decorative: conditional access usually trusts it. If “compliant” can mean “never evaluated”, then every access decision built on it inherits the gap.

First fix: set the built-in default for devices with no policy to not compliant, then assign a baseline policy to all devices with real checks: encryption, minimum OS version, antivirus health. Expect a wave of newly honest non-compliance, and treat it as the finding, not the failure.

3. Conditional access gaps around legacy auth and unmanaged devices

Conditional access estates tend to grow policy by policy, each one solving the incident of the day. What gets missed is the negative space: legacy authentication protocols that were never explicitly blocked, unmanaged personal devices with an unhandled path to corporate data, and break-glass exclusion groups that have swollen into permanent VIP lists.

Legacy protocols cannot enforce multi-factor authentication, which is why password-spray attacks seek them out. And an unmanaged-device gap undoes the work of findings one and two: it does not matter how well managed devices behave if unmanaged ones get the same access.

First fix: use sign-in logs and report-only policies to find what would break, then block legacy authentication outright, require a compliant device for access to corporate data, and audit every exclusion group with a named owner and an expiry.

4. Application deployments stuck at pending

Open the install status report for required applications in almost any tenant and you will find deployments sitting at pending or failed for months, with nobody assigned to ask why. The causes are mundane: detection rules that never match, so an app installs successfully but reports as not installed forever; dependency chains that stall; devices that missed the deployment window and never retried in a way anyone noticed.

The risk depends on the app. A stuck line-of-business tool generates service desk tickets. A stuck security agent generates an audit finding, or worse, an incident on a device everyone believed was protected.

First fix: triage required deployments by failure count, starting with security tooling. Check detection rules before anything else; in our experience a large share of “stuck” apps are installed and healthy, and only the reporting is broken.

5. Autopilot profiles drifted from the signed-off build

The migration project tested a build, documented it and signed it off. Then operations began: an app added to the Enrollment Status Page to fix one complaint, a profile setting toggled to work around one incident, a group tag repurposed for a new office. Each change was reasonable; none went through change control; and two years on, a new starter receives a device no one has ever tested as a whole.

The symptoms are longer provisioning times, intermittent Enrollment Status Page timeouts, and a first-day experience that varies by which change happened to land last.

First fix: provision a reference device today and compare it, setting by setting and app by app, against the signed-off specification. Decide deliberately what the current baseline should be, document it, and put the Autopilot profile and Enrollment Status Page under the same change control as any production system.

The common thread

None of these findings is exotic, and none needs a product to fix. They share one root cause: configuration made under project pressure, operated afterwards without review. A health check is simply the discipline of comparing what the tenant was meant to do with what it is doing now. Our discovery tooling gives us the evidence quickly, but the habit matters more than the tooling: schedule the comparison, or the drift schedules itself.

When did you last check what Intune is actually doing?

Our Intune health check puts evidence behind every one of these findings, from update coverage to Autopilot drift, in weeks not months.