Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Fourward Team: 2 UX Researchers and 2 Product Designers Microsoft XSD Team: 3 Researchers

Role

Design Lead

Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Fourward Team: 2 UX Researchers and 2 Product Designers Microsoft XSD Team: 3 Researchers

Role

Design Lead

Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Fourward Team: 2 UX Researchers and 2 Product Designers Microsoft XSD Team: 3 Researchers

Role

Design Lead

/ context & potential impact

/ context & potential impact

/ context & potential impact

Exploring trust and verification in AI video conferencing tools

When Microsoft Copilot generates meeting summaries, action items, and transcripts, most users skim or skip verification entirely. Inaccurate records become organizational memory. Tasks get assigned that were never agreed upon. Leaders make strategic decisions from incomplete summaries.

Working with Microsoft's XSD team over 6 months, I led research and design for Fourward — a four-person team investigating how knowledge workers decide when to trust and verify AI-generated meeting outputs. As design lead, I drove the research synthesis, facilitated the co-design workshop, made the call to expand our scope from post-meeting to the full meeting lifecycle, and narrowed 20+ co-design concepts into two prototype directions.

Traditional AI meeting workflows place verification at the highest-friction moment — after context has faded, before the next meeting starts. Teams Up moves verification into the meeting itself, producing a recap that arrives already reviewed.

320M+

Microsoft Teams daily users potentially benefiting from Teams Up

1M+

organizations that use Microsoft Teams potentially benefiting from Teams Up

/ The Problem

/ The Problem

AI meeting tools are built for trust. Nobody designed for verification.

Microsoft Copilot and similar tools present all meeting outputs — summaries, action items, decisions — with equal confidence, whether they're accurate or fabricated. Our research identified two compounding gaps in how users interact with these outputs.

Verification Gap 1

Navigation friction

Moving from a claim in the summary to its source in the transcript is slow and disorienting. UTC timestamps, fragmented tool ecosystems, and no inline links mean most verification stops at memory check—the least reliable layer.

Design response → Source tracing

Make the path from claim to source fast enough that verification becomes a natural part of reading, not a separate task.

Verification Gap 2

Fabrication detection

AI confidently asserts content that was never said. Because all outputs look identical regardless of confidence, users have no signal to know what needs checking. Subtle fabrications pass undetected.

Design response → Low-confidence highlighting

Surface AI confidence levels so users know what to verify—without making every output feel uncertain.

Microsoft Teams CoPilot generates AI summaries after user's meetings. These artifacts often go unverified, or very lightly skimmed.

This affects Microsoft's business goal of increasing AI adoption.

56%

power users are more likely to use AI to catch up on missed meetings

44%

lapsed Copilot users cite distrust of answers as the primary reason for stopping use.

/ Key Findings

/ Key Findings

Trust breakdown leads to tool abandonment. Verification should be strategically inserted between these two states.

Finding 01

Users trust the brand before trusting the output, so they skip verification.

"No worries! It's Microsoft CoPilot! It usually gets these things right."

— P5, PM at Big Tech Company

Users skip verification not because they trust the AI's accuracy, but because they trust Microsoft as an institution.

Design response

Surface uncertainty where it exists. Users should be able to trust the parts that are reliable, and quickly identify the parts that need checking.

Finding 02

Verification happens only when four conditions align.

Prior knowledge

Users verify when they have something to check against — they were in the meeting and remember what was said.

Speed

Verification must be fast — under 30 seconds. Anything slower gets abandoned. Navigation friction is the primary cause of verification failure.

High stakes

External sharing, client meetings, and consequential decisions prompt verification. Internal or low-stakes outputs get skimmed or skipped entirely.

Personal responsibility

Users verify when they feel personally accountable for the output being sent. When accountability is diffuse or shared, verification erodes.

Finding 03

Imposed AI creates worker's resistance before the tool is ever used.

When AI tools are introduced institutionally rather than adopted voluntarily, workers develop affective responses (disgust, guilt, ambivalence) that no interface improvement can fix.

"The guy driving the fancy BMW wants us to use AI, but he doesn't even know what for."

— P4, Senior Designer

"You cannot turn it off right now. It's so annoying."

— P5, PM at Big Tech Company

Design response

To correspond with Microsoft's business goal of AI adoption, I delivered a worker accountability and protection framework around AI-generated workflow. Workers feel safer using AI when they know their rights and work are protected.

Microsoft can share this proposed framework with their customers (organizations) and investigate if these policies lead to higher trust and therefore higher organizational AI adoption.

/ full design

/ full design

Live Recap Panel

Overview: Live Meeting Recap Panel. Collapsed and expanded states.

Item cards — editable, traceable, flaggable

Each AI-captured item is a card with four controls: timestamp chip (click to trace to transcript source), flag icon (participants flag for host review), edit icon (host opens inline text field), and delete. Nothing is final until the host confirms it.

Decision cards

Flag for review · Source chip · Inline edit

To do cards

Assign to participants in real time · Three states: assigning / assigned / unassigned

Little Guy — uncertainty through content design

When Microsoft XSD told us confidence scores "communicate that our AI is incapable," we reframed the problem. Instead of warning icons or percentages, the AI uses hedging language embodied by a personified mascot — present in the panel header, animated, communicating status through conversation rather than metrics.

THE PIVOT

Hedging language carries the verification invitation without framing AI as unreliable. Concept validation found participants responded more positively to AI that asked for input than AI that asserted outputs.

Little Guy — character variants

Animated blob mascot — floats in the panel header, cycles through "Listening in…" and "Jotting notes…"

Little Guy — uncertainty nudge

Asks rather than asserts — hedging language replaces confidence scores

End-of-Meeting Co-Review

The host-triggered wrap-up modal is the primary verification moment: two-column layout, all items inline-editable, send immediately or block calendar time for later.

Scrollable recap with summary, decisions, and to-dos

Scrollable recap with Summary, Decisions, and To-Dos. Actions mechanisms are similar to the Live Meeting Recap panel, with the additional Add Item option available to meeting host.

Send recap panel

Send recap panel. Host can select recap participants to send immediately, or choose a time from their calendar to come back and review the recap from their meeting hub.

Post-Meeting Recap Review

Full Recap tab — low-confidence items surfaced

Three content sections: Summary, Decisions, To-Dos. Items the AI couldn't clearly capture are highlighted in orange. A toggle in the AI banner switches between showing and hiding flagged items. The right-hand panel shows the recording and transcript — always available, collapsible to 48px.

Meeting Recap Review. Confidence signaling + source tracing working together.

Inline suggestion cards

Clicking a highlighted row opens a suggestion card beneath it — Little Guy's specific uncertainty question, quick-select option chips for the most likely answers, and a free-text clarification field. Pressing Enter or selecting a chip updates the item text inline and closes the card.

Low-confidence row — editing

Users can click any highlighted row to open the inline suggestion card.

Suggestion card

Little Guy asks a targeted question · Option chips for quick resolution

/ usability tests and the impact

/ usability tests and the impact

The workflow reframe

Three semi-structured think-aloud sessions across the full meeting lifecycle. Participants were framed as hosts responsible for sending an accurate summary while facing immediate time pressure — a back-to-back schedule.

Traditional ai meeting workflow
  • Meeting happens

  • AI generates summary (unreviewed)

  • Verification — if it happens at all

  • Alignment — if context hasn't faded

Verification placed at the highest-friction moment. Low trust leads directly to abandonment.

Teams Up workflow
  • Meeting + Live Recap Panel captures in real time

  • End-of-meeting co-review while memory is live

  • Already-verified recap sent or scheduled

Teams Up delays tool abandonment by inserting a low-cost verification step between low trust and disengagement — positioning the design as an adoption support mechanism, not just a verification aid.

Workflow reframe: traditional four-step flow vs. Teams Up three-step verified-summary flow.

Teams Up delaying tool abandonment by inserting verification between low trust and disengagement.

"I like these real-time stuff! I won't have to take notes and it's easy to use!"

— P1, Usability Study, on the Live Recap Panel

"This is really good! I like [that the AI is] looking over and making sure everything is correct! I'd totally use this!"

— P1, Usability Study, on uncertainty flagging and the recap dashboard

/ Beyond the interface — Trustworthy AI Adoption Framework

/ Beyond the interface — Trustworthy AI Adoption Framework

A strategic output for Microsoft and beyond

Trustworthy AI cannot be built through product design alone. We developed an exploratory Plan → Design → Measure framework — delivered to Microsoft XSD as part of the final handoff — for cross-functional teams thinking about AI adoption at every level.

plan

Why adopt, and who's responsible when it goes wrong?

Define purpose, stage rollout so trust is earned rather than assumed, close the accountability gap by naming who owns AI outputs.

Product leaders, TPMs, customer success

design

How to build trust without adding cognitive burden?

Transparency should be lightweight and contextual. Human control over consequential actions is the mechanism through which trust becomes calibrated.

Product designers, researchers, engineers

measure

Is adoption actually working — beyond usage numbers?

High usage can hide falling trust. Verification time that drops may signal over-reliance. Emotional signals are absent from standard product analytics.

PMs, researchers, customer success

/ What I'm Learning

/ What I'm Learning

Some things I've learned along this project

🧐 Research depth changes what you design.

The difference between designing for a brief and designing from evidence is visible in every decision: why the panel is passive, why it doesn't alert, why the host triggers the review.

🤖 Designing for AI requires designing for trust, not just usability.

Standard usability heuristics — efficiency, learnability, error recovery — don't fully account for the trust dynamics in AI systems. A feature can be perfectly usable and still erode appropriate reliance. We're designing for calibration, not just interaction.

🙆🏻‍♀️ Co-design is a great design method, in addition to being a research method!

The co-design workshop generated 20+ ideas in 2 hours that we wouldn't have reached through desk research alone. More importantly, the workshop revealed which problems participants cared about — the before/during/after expansion came directly from what participants wanted to design for.

💡 Speculative work requires more rigor…

Because we're designing for a future AI capability (confidence tiering in Copilot), every design decision needs stronger justification than product work against existing systems. We can't point to a shipped feature — we have to point to research.

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly