Microsoft 365 monitoring from where your users actually work

Microsoft 365 monitoring measures what your users actually get from Teams, Outlook, SharePoint and OneDrive, from their own device, Wi-Fi, ISP and VPN. Service health reports the incidents Microsoft is working on for your tenant. Experience monitoring adds the view from each desk and home office, and alerts your team when the experience degrades.

Example dashboard: experience scores out of 100 for six Microsoft 365 services; Teams is at 74 and Outlook on the web at 79, the others between 88 and 95. Teams works normally on Windows and macOS but is degraded on the web. Affected users: 112 home workers on ISP B, 38 at the head office, none at the branch office.

Microsoft 365 experience

Score out of 100

Outlook, desktop
92
Teams
74
SharePoint
88
OneDrive
90
Outlook on the web
79
Sign-in
95

Teams, by platform

  • Windows OK
  • macOS OK
  • Web Degraded

Who is affected

  • Head office 38 users
  • Home workers, ISP B 112 users
  • Branch office 0 users

Example data, for illustration only

Two methods

Synthetic tests or real user monitoring: what does each one measure?

Synthetic tests and real user experience monitoring answer different questions. Most organisations use both, in a mix that depends on their sites, devices and priorities.

Compared onSynthetic transactionsReal user experience
How it works Probes repeat scripted actions around the clock from the locations you choose: signing in, sending and receiving mail, posting in Teams, reaching the Teams media relay, opening a file in SharePoint or OneDrive. A light agent on each device measures the device, the local network, the ISP, the VPN and how quickly Microsoft 365 apps respond while people work, and follows each Teams call while it runs.
What it is good at Spotting an outage at night or at the weekend, comparing results over time, and testing hybrid components such as AD FS and Exchange. Showing what each user gets, at home or in the office, by app and platform, against that user's own baseline.
Its limits It does not see the user's own device or home connection. It only covers the devices where the agent is deployed.
Best for Service availability, hybrid dependencies, the path to Teams and sites where nobody works at night. Helpdesk triage, home and hybrid workers, and live Teams call troubleshooting.

We combine both where it helps, with tools chosen for your estate.

Part A

Microsoft 365 monitoring

Six areas we watch for you, from service availability to reports. For each one: what we cover, and how we deliver it with you.

Service availability and outage impact

Know who an outage affects, where and for how long.

When Outlook or Teams slows down, the helpdesk first needs to know whether the cause is Microsoft or something inside your own network. Service health lists the incidents Microsoft is working on for your tenant. Monitoring adds which users and sites are affected and for how long, and keeps a record for SLA follow-up.

What it covers

  • Availability of each Microsoft 365 service, by site
  • Users and locations affected by an incident
  • Microsoft incident or internal issue
  • Duration and impact, kept for SLA follow-up
  • Post-incident report

How we deliver it

  • We set up tests from the sites that matter to you.
  • We agree alert thresholds with your team.
  • Alerts reach the people who can act on them.
  • After a major incident, a report shows who was affected, where and for how long.

Incident impact

Exchange Online incident

Users affected
214
Sites affected
3 of 12
Duration
1 h 40 min
Example data, for illustration only

Teams calls and meetings quality

Troubleshoot a Teams call while it is still running.

When someone says a call is choppy, the cause can be their laptop, their headset, their Wi-Fi, their ISP or the VPN. In the Teams admin center, meetings and events can be examined while they run, and one-to-one and PSTN calls once they have ended. Microsoft Teams monitoring follows every call while it runs, one-to-one, group and PSTN calls included, on Windows and macOS, with fresh figures every 30 seconds. Device load, headset, Wi-Fi, ISP and VPN are recorded with the call, and the history is kept for up to six months.

What it covers

  • Every call followed live: one-to-one, group, PSTN and meetings
  • Audio, video and screen sharing, refreshed every 30 seconds
  • Windows and macOS, with the same detail
  • Device load, headset, Wi-Fi, ISP and VPN recorded with each call
  • Alerts when call quality drops
  • Call history kept for up to six months

How we deliver it

  • We measure call quality where your users sit, at home or in the office.
  • Your helpdesk reads a call while the user is still on the line.
  • We trace poor calls to their likely cause and give your admins a clear fix list.
  • We review the trend with you during the project, then a customer success manager follows up.

Jitter, one call in progress

Live, refreshed every 30 s

Wi-Fi drop

Jitter, one call in progress
10:40 8 ms
10:41 9 ms
10:42 7 ms
10:43 10 ms
10:44 9 ms
10:45 11 ms
10:46 14 ms
10:47 22 ms
10:48 37 ms
10:49 46 ms
Example data, for illustration only

Teams media path and relays

Check the path every Teams call would take, before anyone calls.

Teams sends audio and video straight to the other party when it can, and through a Microsoft media relay when it cannot. Microsoft recommends keeping Teams media out of the VPN tunnel. When DNS, the VPN or a proxy sends a user to a distant relay, every call they make carries the delay. Synthetic tests on each monitored device check which relay a call would use right now and how long it takes to reach it, measure every hop on the way, and confirm that the VPN split tunnel works as intended for Teams.

What it covers

  • The Teams media relay each device would use, checked continuously
  • Latency to the relay, device by device and site by site
  • Distant relays caused by DNS, VPN or proxy settings
  • VPN split tunnel checked for Teams traffic
  • Every hop from the device to Microsoft's network, with 90 days of history
  • Alerts when the path to Teams slows down

How we deliver it

  • We map your sites and remote users to the relays they reach.
  • We point out the DNS, VPN or proxy settings that send calls the long way.
  • Your network team gets the hop at fault, with figures to share with the ISP.
  • We review the results with you during the project, then a customer success manager follows up.

Teams media path, one user

Checked continuously

  1. Laptop Normal
  2. Home Wi-Fi 9 ms
  3. ISP 14 ms
  4. VPN Split tunnel OK
  5. Teams media relay West US, 142 ms
  6. Microsoft network Available
Example data, for illustration only

User experience by app and platform

See what each user gets, app by app and platform by platform.

The same service can work well in Outlook on the desktop and badly in Outlook on the web, or slow down only for the customers of one ISP. End user experience monitoring measures how Outlook, Teams, SharePoint and OneDrive respond on each platform, and compares the result with each user's own baseline, so a real degradation stands out from normal variation.

What it covers

  • Outlook, Teams, SharePoint and OneDrive
  • Desktop, web and mobile apps
  • Windows and macOS devices
  • Experience score against each user's own baseline
  • Device health under load: CPU and memory

How we deliver it

  • We agree with you which apps and user groups to start with.
  • The agent is rolled out in waves, starting with a pilot group.
  • Your helpdesk learns to read a user's history before calling back.
  • We review the scores with you and tune the thresholds.

Experience by app and platform

Against each user's baseline

  • Outlook, desktop Normal
  • Outlook on the web Slower than usual
  • Teams, macOS Normal
  • SharePoint, ISP B Degraded
Example data, for illustration only

Network and application monitoring

Find the step on the network path that slows your apps down.

Between a user and Microsoft 365 sit the device, the Wi-Fi, the ISP, the VPN and the corporate network. A slowdown on any of them looks like a Microsoft problem to the user. Monitoring measures each step, ranks ISPs and managed networks by the experience they deliver, and runs synthetic tests on key applications, so you see where the time goes.

What it covers

  • Every step: device, Wi-Fi, ISP, VPN, corporate network, Microsoft cloud
  • ISPs and managed networks ranked by experience
  • Synthetic tests on key applications
  • Alerts that point to the step at fault
  • History for each incident review

How we deliver it

  • We choose the tests with you, app by app.
  • Tests run around the clock, nights and weekends included.
  • Each alert points to the step at fault.
  • You keep the history for every incident review.

Network path, one user

Opening a SharePoint page

  1. Device Normal
  2. Home Wi-Fi 62 ms
  3. ISP 14 ms
  4. VPN 21 ms
  5. Corporate network 8 ms
  6. Microsoft 365 Available
Example data, for illustration only

Hybrid dependencies

Know when an on-premises component puts sign-in or mail at risk.

Many tenants still depend on servers you run yourself: AD FS or pass-through authentication for sign-in, Exchange hybrid for mail and calendars, and the certificates behind them. When one of them fails or a certificate expires, users can lose sign-in, mail flow or calendar sharing even though Microsoft's services are healthy. Synthetic tests check these components, and certificate expiry raises an alert well before the date.

What it covers

  • AD FS sign-in
  • Pass-through authentication
  • Exchange hybrid mail flow and connectors
  • Certificate expiry dates

How we deliver it

  • We list your hybrid dependencies with your team before we start.
  • Each one gets a synthetic test and an owner.
  • Certificate alerts go out well before expiry.
  • Exchange Server itself is covered in part B.

Hybrid checks

3 of 4 passing

  • AD FS sign-in
  • Pass-through authentication
  • Mail flow to Exchange Online
  • Certificate expires in 21 days
Example data, for illustration only

Reporting and analytics

Reports for each audience, from the helpdesk to management.

Raw monitoring data rarely answers the question a manager asks. Dashboards are set up for each team: the helpdesk, the unified communications team, the NOC and management. Built-in reports cover licence use, Teams activity and mail traffic, and reports set up with you explain in plain language what changed and why.

What it covers

  • Dashboards for the helpdesk, UC team, NOC and management
  • Availability and SLA follow-up
  • Licence use, Teams activity and mail traffic
  • Trends by site, ISP and service
  • Reports set up with you, in plain language

How we deliver it

  • We agree the indicators that matter to you.
  • Reports arrive on schedule, in plain language.
  • We highlight what changed and why.
  • During the project, we help you rank the next fixes by impact.

Experience score by ISP

  • ISP A: This month 86, Baseline 84
  • ISP B: This month 61, Baseline 79
  • ISP C: This month 80, Baseline 81
  • Managed: This month 90, Baseline 88
Example data, for illustration only

Deployment

How is experience monitoring deployed?

The tools we put in place roll out with what you already use. We confirm the details for your estate before we start.

  • An agent on each device

    A light agent runs in the user's context, without admin rights, on Windows and macOS devices and on virtual desktops such as Citrix and Azure Virtual Desktop.

  • Your usual deployment tools

    It is rolled out with Microsoft Intune, Configuration Manager, Group Policy or any other software deployment tool.

  • Probes for synthetic tests

    Synthetic tests run from probes installed on workstations, at the sites you choose.

  • Permissions reviewed first

    Reports that read Microsoft 365 data use an app in Microsoft Entra ID, with permissions your security team reviews before anything runs.

How it works

How monitoring runs for your users

Fast and done right: ready-made tooling, tested starting thresholds and runbooks your team keeps.

  1. Step 1

    Scope

    We agree with you which apps, sites and user groups to watch first, and who receives which alert.

  2. Step 2

    Deploy

    Probes and agents are rolled out in waves, starting with a pilot group.

  3. Step 3

    Baseline

    Each user's normal experience is measured first, so alerts compare like with like.

  4. Step 4

    Alert

    Thresholds and recipients are agreed with your helpdesk, UC, network and NOC teams.

  5. Step 5

    Report

    Dashboards for each team, and reports set up with you for IT and management.

FAQ

Monitoring questions, answered.

What is Microsoft 365 monitoring?

It is the continuous measurement of how Microsoft 365 performs for your users: whether services are available, how quickly Outlook, Teams, SharePoint and OneDrive respond, and what the network path between the user and Microsoft adds. It combines synthetic tests with measurements taken on each user's device, and alerts the right team when the experience degrades. It is also called M365 or Office 365 monitoring.

What does Service health show, and what does experience monitoring add?

Service health, in the Microsoft 365 admin center, lists the incidents and advisories Microsoft is working on for your tenant, service by service. Experience monitoring adds what each user actually gets, measured from their own device, Wi-Fi, ISP and VPN, so you can tell whether a slowdown comes from the service, the network or the device. We use both.

What is digital employee experience (DEX) monitoring?

Digital employee experience monitoring measures how well the tools people use every day work for them, from the device to the application. For Microsoft 365, that means app response times, call quality, device health and network quality, measured continuously rather than only when someone complains.

Synthetic or real user monitoring: which do we need?

Usually both. Synthetic tests catch an outage at night and test hybrid components such as AD FS and Exchange. Real user monitoring shows what each person gets, at home or in the office. We look at your sites, devices and priorities, then propose the mix and the tools.

Does it work for home and hybrid workers?

Yes. The agent measures from the user's own device, so home Wi-Fi, the home ISP and the VPN are part of the picture. Scores are compared with each user's own baseline, so a remote worker's bad day stands out.

Can you troubleshoot a single Teams call?

Yes, while it is still running. The Teams admin center shows meetings and events live, and one-to-one and PSTN calls once they have ended, after processing that can take from 30 minutes to two hours. Experience monitoring follows every call as it happens, on Windows and macOS, with the device, headset, Wi-Fi, ISP and VPN recorded alongside, so the helpdesk can act while the user is still on the line. The history is kept for later reviews.

Why do some Teams calls go through a distant relay?

Teams sends audio and video straight to the other party when it can, and through a Microsoft relay when it cannot. Microsoft recommends keeping Teams media out of the VPN tunnel. When DNS, the VPN or a proxy sends a device to a distant relay, every call it makes carries the delay. From each monitored device, we check which relay a call would use right now and how long it takes to reach it, then point to the setting at fault.

Who receives the alerts?

The people who can act on them. We agree thresholds and recipients with your team: the helpdesk, the unified communications team, the network team or the NOC. Each alert points to the likely cause, such as the service, a step on the network path or the device.

Talk to us about your users' experience.

Tell us where your users feel the slowdowns. We will explain which tools fit your sites and devices, and how we would put them in place with your team.