Exploring trust and verification in AI video conferencing tools
When Microsoft Copilot generates meeting summaries, action items, and transcripts, most users skim or skip verification entirely. Inaccurate records become organizational memory. Tasks get assigned that were never agreed upon. Leaders make strategic decisions from incomplete summaries.
Working with Microsoft's XSD team over 6 months, I led research and design for Fourward — a four-person team investigating how knowledge workers decide when to trust and verify AI-generated meeting outputs. As design lead, I drove the research synthesis, facilitated the co-design workshop, made the call to expand our scope from post-meeting to the full meeting lifecycle, and narrowed 20+ co-design concepts into two prototype directions.
Traditional AI meeting workflows place verification at the highest-friction moment — after context has faded, before the next meeting starts. Teams Up moves verification into the meeting itself, producing a recap that arrives already reviewed.
320M+
Microsoft Teams daily users potentially benefiting from Teams Up
1M+
organizations that use Microsoft Teams potentially benefiting from Teams Up
AI meeting tools are built for trust. Nobody designed for verification.
Microsoft Copilot and similar tools present all meeting outputs — summaries, action items, decisions — with equal confidence, whether they're accurate or fabricated. Our research identified two compounding gaps in how users interact with these outputs.
Verification Gap 1
Navigation friction
Moving from a claim in the summary to its source in the transcript is slow and disorienting. UTC timestamps, fragmented tool ecosystems, and no inline links mean most verification stops at memory check—the least reliable layer.
Design response → Source tracing
Make the path from claim to source fast enough that verification becomes a natural part of reading, not a separate task.
Verification Gap 2
Fabrication detection
AI confidently asserts content that was never said. Because all outputs look identical regardless of confidence, users have no signal to know what needs checking. Subtle fabrications pass undetected.
Design response → Low-confidence highlighting
Surface AI confidence levels so users know what to verify—without making every output feel uncertain.
Microsoft Teams CoPilot generates AI summaries after user's meetings. These artifacts often go unverified, or very lightly skimmed.
This affects Microsoft's business goal of increasing AI adoption.
56%
power users are more likely to use AI to catch up on missed meetings
44%
lapsed Copilot users cite distrust of answers as the primary reason for stopping use.
Trust breakdown leads to tool abandonment. Verification should be strategically inserted between these two states.
Finding 01
Users trust the brand before trusting the output, so they skip verification.
"No worries! It's Microsoft CoPilot! It usually gets these things right."
— P5, PM at Big Tech Company
Users skip verification not because they trust the AI's accuracy, but because they trust Microsoft as an institution.
Design response
Surface uncertainty where it exists. Users should be able to trust the parts that are reliable, and quickly identify the parts that need checking.
Finding 02
Verification happens only when four conditions align.
Prior knowledge
Users verify when they have something to check against — they were in the meeting and remember what was said.
Speed
Verification must be fast — under 30 seconds. Anything slower gets abandoned. Navigation friction is the primary cause of verification failure.
High stakes
External sharing, client meetings, and consequential decisions prompt verification. Internal or low-stakes outputs get skimmed or skipped entirely.
Personal responsibility
Users verify when they feel personally accountable for the output being sent. When accountability is diffuse or shared, verification erodes.
Finding 03
Imposed AI creates worker's resistance before the tool is ever used.
When AI tools are introduced institutionally rather than adopted voluntarily, workers develop affective responses (disgust, guilt, ambivalence) that no interface improvement can fix.
"The guy driving the fancy BMW wants us to use AI, but he doesn't even know what for."
— P4, Senior Designer
"You cannot turn it off right now. It's so annoying."
— P5, PM at Big Tech Company
Design response
To correspond with Microsoft's business goal of AI adoption, I delivered a worker accountability and protection framework around AI-generated workflow. Workers feel safer using AI when they know their rights and work are protected.
Microsoft can share this proposed framework with their customers (organizations) and investigate if these policies lead to higher trust and therefore higher organizational AI adoption.
Live Recap Panel
Overview: Live Meeting Recap Panel. Collapsed and expanded states.
Item cards — editable, traceable, flaggable
Each AI-captured item is a card with four controls: timestamp chip (click to trace to transcript source), flag icon (participants flag for host review), edit icon (host opens inline text field), and delete. Nothing is final until the host confirms it.
Decision cards
Flag for review · Source chip · Inline edit

To do cards
Assign to participants in real time · Three states: assigning / assigned / unassigned

Little Guy — uncertainty through content design
When Microsoft XSD told us confidence scores "communicate that our AI is incapable," we reframed the problem. Instead of warning icons or percentages, the AI uses hedging language embodied by a personified mascot — present in the panel header, animated, communicating status through conversation rather than metrics.
THE PIVOT
Hedging language carries the verification invitation without framing AI as unreliable. Concept validation found participants responded more positively to AI that asked for input than AI that asserted outputs.
Little Guy — character variants
Animated blob mascot — floats in the panel header, cycles through "Listening in…" and "Jotting notes…"

Little Guy — uncertainty nudge
Asks rather than asserts — hedging language replaces confidence scores



End-of-Meeting Co-Review
The host-triggered wrap-up modal is the primary verification moment: two-column layout, all items inline-editable, send immediately or block calendar time for later.
Scrollable recap with summary, decisions, and to-dos

Scrollable recap with Summary, Decisions, and To-Dos. Actions mechanisms are similar to the Live Meeting Recap panel, with the additional Add Item option available to meeting host.
Send recap panel

Send recap panel. Host can select recap participants to send immediately, or choose a time from their calendar to come back and review the recap from their meeting hub.
Post-Meeting Recap Review
Full Recap tab — low-confidence items surfaced
Three content sections: Summary, Decisions, To-Dos. Items the AI couldn't clearly capture are highlighted in orange. A toggle in the AI banner switches between showing and hiding flagged items. The right-hand panel shows the recording and transcript — always available, collapsible to 48px.
Meeting Recap Review. Confidence signaling + source tracing working together.
Inline suggestion cards
Clicking a highlighted row opens a suggestion card beneath it — Little Guy's specific uncertainty question, quick-select option chips for the most likely answers, and a free-text clarification field. Pressing Enter or selecting a chip updates the item text inline and closes the card.
Low-confidence row — editing
Users can click any highlighted row to open the inline suggestion card.

Suggestion card
Little Guy asks a targeted question · Option chips for quick resolution

The workflow reframe
Three semi-structured think-aloud sessions across the full meeting lifecycle. Participants were framed as hosts responsible for sending an accurate summary while facing immediate time pressure — a back-to-back schedule.
Traditional ai meeting workflow
Meeting happens
AI generates summary (unreviewed)
Verification — if it happens at all
Alignment — if context hasn't faded
Verification placed at the highest-friction moment. Low trust leads directly to abandonment.
Teams Up workflow
Meeting + Live Recap Panel captures in real time
End-of-meeting co-review while memory is live
Already-verified recap sent or scheduled
Teams Up delays tool abandonment by inserting a low-cost verification step between low trust and disengagement — positioning the design as an adoption support mechanism, not just a verification aid.

Workflow reframe: traditional four-step flow vs. Teams Up three-step verified-summary flow.

Teams Up delaying tool abandonment by inserting verification between low trust and disengagement.
"I like these real-time stuff! I won't have to take notes and it's easy to use!"
— P1, Usability Study, on the Live Recap Panel
"This is really good! I like [that the AI is] looking over and making sure everything is correct! I'd totally use this!"
— P1, Usability Study, on uncertainty flagging and the recap dashboard
A strategic output for Microsoft and beyond
Trustworthy AI cannot be built through product design alone. We developed an exploratory Plan → Design → Measure framework — delivered to Microsoft XSD as part of the final handoff — for cross-functional teams thinking about AI adoption at every level.
plan
Why adopt, and who's responsible when it goes wrong?
Define purpose, stage rollout so trust is earned rather than assumed, close the accountability gap by naming who owns AI outputs.
Product leaders, TPMs, customer success
design
How to build trust without adding cognitive burden?
Transparency should be lightweight and contextual. Human control over consequential actions is the mechanism through which trust becomes calibrated.
Product designers, researchers, engineers
measure
Is adoption actually working — beyond usage numbers?
High usage can hide falling trust. Verification time that drops may signal over-reliance. Emotional signals are absent from standard product analytics.
PMs, researchers, customer success
Some things I've learned along this project
🧐 Research depth changes what you design.
The difference between designing for a brief and designing from evidence is visible in every decision: why the panel is passive, why it doesn't alert, why the host triggers the review.
🤖 Designing for AI requires designing for trust, not just usability.
Standard usability heuristics — efficiency, learnability, error recovery — don't fully account for the trust dynamics in AI systems. A feature can be perfectly usable and still erode appropriate reliance. We're designing for calibration, not just interaction.
🙆🏻♀️ Co-design is a great design method, in addition to being a research method!
The co-design workshop generated 20+ ideas in 2 hours that we wouldn't have reached through desk research alone. More importantly, the workshop revealed which problems participants cared about — the before/during/after expansion came directly from what participants wanted to design for.
💡 Speculative work requires more rigor…
Because we're designing for a future AI capability (confidence tiering in Copilot), every design decision needs stronger justification than product work against existing systems. We can't point to a shipped feature — we have to point to research.
2025
Eliminate context-switching disruptions for healthcare hiring managers' credential verification workflow
/ vendor management system / design system
Sony Electronics
2025
Improve the accessibility and OOBE of Sony LinkBuds Open earbuds for users with visual impairments and dexterity challenges
/ accessibility evaluation

