Microsoft · 2024–25

The question

How can Copilot earn users’ trust?

Microsoft had conducted several studies about Copilot and AI use across Outlook, but the findings had not been synthesized into a shared product perspective. I reviewed a year of internal research, public conversations, and academic literature to develop a white paper on user trust and product priorities.

12+ internal studies Point-of-view white paper Under NDA

My role

UX Researchercontract, via Kadence

Timeline

Dec 2024 – Feb 2025paper delivered Feb 2025

Inputs

A year of internal studiesplus public social listening and academic literature

Deliverable

A point-of-view white paperwith design guidelines and a feature checklist

Why it mattered

Outlook teams needed a shared research perspective on AI integration.

Copilot was expanding across Microsoft 365, while research remained organized around individual features. Teams had separately studied drafting, inbox prioritization, scheduling, summaries, and automation. Across these studies, similar questions emerged about trust, user control, proactive assistance, and automation. I synthesized the findings into a shared point of view on user needs and product priorities for Copilot in Outlook.

What was needed

How should AI enter the email ecosystem without breaking trust?

A point of view on product priorities, the information and controls users need, and the requirements for more autonomous assistance.

Method

How I approached the synthesis

The synthesis at a glance

Inputs

Primary

12+ internal studies

drafting, prioritization, time-away, concept focus groups, workflow, out-of-office

Validation

Public social listening

used to assess themes identified in internal research

Grounding

Prior literature & internal frameworks

Analysis

Thematic synthesis and framework development

Outputs

A trust framework

transparency, control, learnability, staged autonomy

A pre-ship checklist

seven checklist categories a team runs before anything ships

Per-feature adoption guidance

criteria for assigning capabilities to each stage of autonomy

1

Research review and extraction

I extracted findings, open questions, participant concerns, and recommendations from each report into a shared repository. Each note retained its original source.

2

Source and study review

For each finding, I reviewed the original study, participant sample, research questions, and strength of support. Findings from small samples or exploratory studies were retained with those limitations documented.

where is this finding coming from? which study, what sample? what does success even look like here? define it before claiming it need to know the current implementation before I'm allowed an opinion
3

Cross-study thematic analysis

I compared findings across studies with different research objectives. Research on inbox prioritization, drafting, out-of-office catch-up, project workflows, and automation identified recurring needs related to workload, transparency, correction, and control.

4

Public conversation review

I compared themes from the internal research with public conversations about AI assistants and Microsoft 365 Copilot. Public posts were used as a secondary check on the synthesis and were not treated as participant research.

X · Microsoft 365

The Teams call summary is a genuine win. This is the kind of thing people actually asked for.

ZDNet

To a lot of users, the 365 Copilot launch read like a mess.

Reddit · r/sysadmin

Half the threads were people asking how to switch it off, not how to use it.

Threads · LinkedIn

Forced opt-in was the fastest way to make people resent a feature that might have helped them.

Paraphrased public posts reviewed during the synthesis.

Synthesis

Copilot as a Chief of Staff

Across the 12 studies, participants described related needs for contextual assistance, support with routine work, and visibility into AI decisions. I organized these findings around a chief-of-staff model, which gave teams a shared way to evaluate capabilities across Outlook.

How it was being built

Prioritize inbox Summarize threads Draft replies Suggest meeting times Track to-dos Search across apps

Each one shipped and measured on its own, by the team that owned it.

What the paper argued for

One assistant that behaves like a chief of staff, working in your flow and on your behalf.

Triages the inbox Catches you up Drafts in your voice Guards your calendar Tracks loose ends Resurfaces what matters

The research also showed a need for continuity across Mail, Calendar, Chat, and Docs. Separate memory in each product would require users to repeatedly provide the same context and preferences. The paper therefore recommended a consistent approach to context across Outlook experiences.

1

On-demand

2

Proactive

3

Automatic

01

On-demand

The user asks; the assistant answers; the user reviews the output before using it.

Examples included summarizing a thread, drafting a reply, and finding an email. Across the studies, participants were most comfortable with this level of assistance.

02

Proactive

The assistant surfaces summaries, suggestions, and priorities without being asked, but every action still belongs to the user.

At this stage, the assistant must decide what is worth surfacing and explain why. Irrelevant or unexplained suggestions could be perceived as noise.

03

Automatic

The assistant acts on its own, sending, archiving, rescheduling, within boundaries the user has set.

The paper recommended moving a task to this stage only after users had experience with its on-demand and proactive versions.

The paper proposed introducing assistance in stages so users could understand and evaluate a capability before it became more autonomous.

The framework

A framework for increasing autonomy

The vision it all sits under

Copilot as a chief of staff

InterfaceTransparency & control

Show what Copilot looked at, why it made a recommendation, and how the user can correct or limit it in the moment.

IntelligenceLearning

Improve from user corrections and carry that learning across sessions, so Copilot does not feel like it resets every time it is opened.

FoundationReliability

Get the basics right first: names, dates, people, deadlines, and context. A confident wrong answer is worse than an incomplete one.

Staged autonomy

Move from on-demand to proactive to automatic assistance one task at a time, based on how users evaluate each workflow.

Across the studies, problems with trust were often associated with limited transparency, control, or learnability.

T

Transparency

"What did it look at, and why did it suggest that?"

Users were more willing to evaluate Copilot's output when they could see the sources and reasoning behind it. Without that visibility, even useful suggestions could feel arbitrary.

C

Control

"Can I correct it, limit it, or undo it here?"

Control had to exist at the moment of use. A setting buried elsewhere did not help when Copilot made a questionable judgment in the inbox.

L

Learnability

"Does it get better after I correct it?"

Users did not expect Copilot to be perfect. They expected it to learn. Repeated mistakes were especially damaging because they made correction feel pointless.

Concept example

Applying the framework to inbox prioritization

The example shows how explanations can help users evaluate AI-generated priorities.

Inbox, prioritized by AI

Explain decisions
Maya Chen Re: Q3 budget review, need your sign-off by EOD High priority

Why: from your direct manager, contains a deadline today, and you replied to this thread twice this week.

Atlas Conference Early-bird tickets end soon! Low priority

Why: bulk promotional sender. You've never opened the previous 14 emails from this address.

Sam Okafor Quick question about the handoff doc High priority

Why: direct question addressed to you, from a close collaborator, on a document you edited yesterday.

Expected user response

"It marked something high priority. I have no idea why, so I'll check everything myself anyway."

Design recommendation: explain why the system assigned each priority.

The deliverable

What the paper gave product teams

The paper documented the research sources, synthesis, framework, design guidelines, and feature checklist for product teams.

The trust & safety checklist

A pre-ship checklist for feature teams

The white paper included a pre-ship checklist for Copilot feature teams. Before making a capability more proactive or automatic, teams could assess whether users understood it, could correct or limit it, and could recover from mistakes.

Transparency4 checks
  • Can someone tell what the feature actually does?
  • Are the data sources it draws on shown?
  • Is the reason behind each decision explained?
  • Can users see the actions it has already taken?
Accuracy & trust4 checks
  • Are the results reliable and consistent?
  • Is AI-generated content clearly marked as such?
  • Does it say so when it is uncertain?
  • Does it avoid presenting invented information as fact?
Learning4 checks
  • Does it learn from the way people correct it?
  • Does that learning survive across sessions?
  • Can users see and edit what it has learned about them?
  • Does it visibly improve over time?
User control4 checks
  • Can users set limits on what it does automatically?
  • Is there an easy way to correct its mistakes?
  • Is the control granular, not all-or-nothing?
  • Can users switch it off or scope it down?
Discoverability3 checks
  • Is the feature easy to find and start using?
  • Are its capabilities communicated clearly?
  • Does it prove its value on first use?
Risk management4 checks
  • Are there safeguards against critical errors?
  • Is there a rollback path when it gets something wrong?
  • Are high-risk actions flagged and confirmed first?
  • Does it fail gracefully and handle sensitive data with care?
Integration4 checks
  • Is it consistent with the other Copilot surfaces?
  • Are a user's preferences respected across features?
  • Does context move with the user instead of resetting per app?
  • Does it present as one assistant, not many?
Where the framework stops

The framework assumes a good-faith user and an honest environment. It does not cover adversarial use: an assistant that sends and replies on someone's behalf is also a new surface for impersonation and phishing, and a system that learns from corrections can be deliberately mistrained. Those questions were outside this paper's scope.

Impact

How the work was used

The paper was shared with relevant product teams, and the checklist was adopted as a review template for Copilot features.