Healthy engagement is repeated behavior that helps customers accomplish a worthwhile job at an appropriate frequency and acceptable effort. Activity without value produces events, sessions, or time spent without reliably producing that outcome. Measure successful use and customer experience together before deciding that more activity is better.
For some products, better engagement means more recurring use. For others, it means accomplishing the same job in fewer visits. Your measurement should reflect what the customer is trying to do rather than reward the interface for keeping them busy.
Start with the customer’s job and cadence
Describe the outcome outside your product vocabulary. An accountant wants an accurate reconciled statement. A hiring manager wants to decide which candidates to advance. A coordinator wants a shift filled. These outcomes are different from opening a dashboard, adding a record, or clicking a notification.
Determine when the need recurs. Daily activity is a poor default for a monthly reporting job. A weekly communication job may need frequent contribution and timely responses. A one-time task can end successfully without any return visit.
Define the eligible population as well. An employee with no task this week should not automatically count as disengaged. Separate lack of opportunity from failure to achieve the outcome when the opportunity exists. Keep the eligibility rule independent of the intervention you want to evaluate wherever possible.
Build an outcome, effort, and experience bundle
Google Research’s HEART paper describes user-centered measurement categories and a process for mapping goals to metrics. It is useful because engagement is considered alongside other aspects of experience rather than treated as a standalone goal.
Our practical bundle has three parts. Outcome measures indicate whether the job was completed correctly. Effort measures indicate what it took. Experience measures indicate whether customers understood and trusted the result. Repeated use adds a time dimension to all three.
| Measurement layer | Reporting product example | Possible misleading substitute |
|---|---|---|
| Successful outcome | Correct report shared before its deadline | Report editor opened |
| Effort | Time and corrections required to finish | Total session duration |
| Experience | Confidence in source accuracy and next steps | A single satisfaction score without context |
| Repetition | Next required report completed | Any return visit |
Do not assume that a click event confirms correctness. A submitted report can still contain errors. Add a suitable quality signal, such as validation failures, later corrections, or review outcomes, and understand the cases your signal misses.
Read activity in both directions
A rise in session count could mean more demand, broader adoption, or a broken workflow requiring repeated attempts. A fall could mean abandonment, successful automation, or fewer opportunities to use the product. The event trend alone does not distinguish these explanations.
Pair activity with the job’s endpoint. If sessions rise while successful tasks remain flat and errors increase, investigate friction. If sessions fall while correct outcomes rise and customer effort decreases, the product may be serving people better.
Microsoft Research’s metric taxonomy separates overall outcomes, diagnostic feature metrics, guardrails, and data-quality metrics. Applying that distinction helps prevent a local activity increase from becoming the sole reason to ship a change.
Worked example: a synthetic expense workflow
A fictional expense product tests an automatic receipt-matching feature. The customer job is to submit an accurate expense claim with minimal effort. All figures below are invented examples.
| Weekly measure per eligible employee | Existing workflow | Automatic matching |
|---|---|---|
| Product sessions | 4 | 2 |
| Receipt-edit actions | 12 | 5 |
| Accurate claims submitted | 0.8 | 0.9 |
| Median minutes per completed claim | 14 | 8 |
| Claims requiring later correction | 6% | 5% |
If you watched sessions and edits alone, the new workflow would appear less engaging. The broader bundle suggests a potentially better outcome with less effort. It still requires a suitable comparison, uncertainty estimates, and enough time to see whether correction rates remain acceptable.
Now imagine another variation raises sessions from four to six but leaves accurate claims at 0.8 and increases later corrections to 9%. That pattern is consistent with activity without improved value. The team should investigate the mechanism rather than celebrate the increase.
A third variation could increase accurate claims while pushing work onto a manager who reviews them. Measure the wider workflow. Moving effort from one role to another is different from removing it.
Define meaningful engagement at the right unit
For a collaboration product, account-level success may require several roles. One person creates a draft, another reviews it, and a third makes a decision. Ten drafts without any review may represent accumulation rather than value.
Define a meaningful account cycle: a relevant artifact created, reviewed, and used. Track the role distribution and timing. Preserve user-level signals so an apparently healthy account does not hide an overloaded administrator or excluded participants.
For an automated product, customers may receive value without opening the interface. A background job that completes correctly and delivers an understood result can matter more than a visit. Ensure that your instrumentation observes the delivery and outcome rather than only the interface.
Use qualitative evidence to interpret the bundle
Ask customers to show a recent successful task and a recent difficult one. Compare their descriptions with the measured events. An analyst may call an export “success” while the customer describes spending another hour fixing it elsewhere.
Sample people who abandoned, people who completed the job, and people with unusual levels of activity. Power users can explain important workflows, but they may tolerate friction that makes the product inaccessible to others.
Use surveys sparingly and at relevant moments. A question about confidence after a completed task is easier to interpret than an unsolicited satisfaction prompt during setup. Keep response selection in mind: the people who answer may differ from the people who leave.
Printable engagement-definition worksheet
| Field | Your definition |
|---|---|
| Customer job and desired result | __________ |
| Who has an opportunity to perform it? | __________ |
| Expected cadence or deadline | __________ |
| Successful outcome event | __________ |
| Quality check or failure signal | __________ |
| Customer effort and displaced work | __________ |
| Experience question and timing | __________ |
| Repeat-success definition | __________ |
| Activity metrics used only as diagnostics | __________ |
| Account and role-level view | __________ |
| What would make more activity a bad sign? | __________ |
| Intervention and evaluation horizon | __________ |
Evaluate a change without rewarding busywork
State the outcome you expect to improve and the activity changes you expect to see. An automation feature may intentionally reduce clicks. A collaboration feature may intentionally increase contributions, but should also improve the completed shared job.
Use stable eligibility and comparable observation windows. Track data quality, because a release can change event collection without changing behavior. If an automatic process replaces a manual event, update the measurement carefully and avoid treating a missing old event as customer loss.
Set guardrails for accuracy, accessibility, latency, workload, and unwanted notifications as appropriate. Select a few that reflect actual plausible harms rather than filling a generic checklist. Review both overall results and role-specific consequences where the intervention affects a team.
Limits of engagement proxies
Customer value can be delayed, subjective, or hard to observe. A correct recommendation might still lead to a poor later decision. A learning product can produce completion without understanding. No single event bundle will fully capture these outcomes.
Treat the bundle as a working approximation and improve it with customer evidence. Do not demand unnecessary events simply to make the metric look healthier. The most useful engagement definition is one that can recognize both more successful recurring use and the welcome disappearance of unnecessary work.
Sources
Something we should correct?
Tell the editors ↗