Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Research Lead: Mai Kana, Ayumi Oishi Design Lead: Han Khuong, Tsai-Ping Kuo Sponsors: Microsoft XSD — Madeline Schenck, Kari Rowles, Sarah Welty

Role

Design Lead & Researcher (Trust Behaviors)

Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Research Lead: Mai Kana, Ayumi Oishi Design Lead: Han Khuong, Tsai-Ping Kuo Sponsors: Microsoft XSD — Madeline Schenck, Kari Rowles, Sarah Welty

Role

Design Lead & Researcher (Trust Behaviors)

Designing for calibrated trust and verification in enterprise agentic AI video conferencing tools

2026

Project Type

Corporate Sponsorship

Duration

January – June 2026

Tools

Figma, Elicit, Claude

Deliverables

2 interactive prototypes, Trustworthy AI Adoption framework

Team

Research Lead: Mai Kana, Ayumi Oishi Design Lead: Han Khuong, Tsai-Ping Kuo Sponsors: Microsoft XSD — Madeline Schenck, Kari Rowles, Sarah Welty

Role

Design Lead & Researcher (Trust Behaviors)

/ my role and contributions

/ my role and contributions

/ my role and contributions

I led a project on trust and verification for agentic AI systems, in collaboration with Microsoft XSD researchers.

My contributions
  • Created user trust behavior framework from synthesizing 30+ research papers on sociology, service design, and automation

  • Designed and piloted the user interview protocol (7 participants, 60 min each)

  • Led co-design workshop facilitation: scenario design, warm-up, synthesis

  • Expanded scope from post-meeting to the full meeting lifecycle

  • Narrowed 20+ co-design concepts to two focused prototype directions

  • Built interactive HTML prototypes used in concept validation and usability testing

Fourward team alongside our instructor during the MS HCDE capstone showcase!

/ Microsoft's business problem

/ Microsoft's business problem

44% of lapsed Copilot users cite distrust as the reason they stopped.

Trust concerns are the primary barrier to production deployment. So the brief we got from the Microsoft XSD team for this project was deliberately open — how do you design for trust and verification in AI?

This pushed me to think about domains. AI is used in multiple tools across the Microsoft 365 Suite. Therefore, my next question was:

In which context do trust and verification matter most for AI adoption?

320M+

active Microsoft Teams daily users

56%

power users are more likely to use AI to catch up on missed meetings

Microsoft Teams CoPilot generates AI summaries after user's meetings. These artifacts often go unverified, or very lightly skimmed.

After researching Microsoft's Work Trend Index, Copilot adoption trends and statistics, we narrowed down to workplace video conferencing. This is a domain with established AI use cases, rapid adoption growth, and distrust as a documented reason for abandonment. Verification behavior was concrete and observable. We could study real practices, not just attitudes.

/ research — understanding the present

/ research — understanding the present

Understanding the relationship between trust and verification in video conferencing

the trust onion

Trust operates across layers, systems, and time. But across all layers, the goal is always "trust calibration."

Trust operates across layers, systems, and time. But across all layers, the goal is always trust calibration.

Literature review synthesis mapped the concept of trust through 4 layers (like an onion!). It operates across all four, with human trust as the core, and each layer after builds upon the foundational philosophical concept of "human trust."

“Good calibration [...] can mitigate misuse and disuse of automation, and so they can guide design, evaluation, and training to enhance human-automation partnerships.”

Lee, J. D., & See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance.

calibrated trust in practice = verification

Trust breakdown leads to tool abandonment. Verification needs to be inserted between these two states.

the verification cascade

Users expressed that their workflow generally consisted of four escalating layers — but most stop at layer one, the most unreliable.

why do they stop at the first layer?
Verification Gap 1

Source tracing friction

"I remember like, hey, I want to refer back to [this point], but I don't even know where it is. Everything is just everywhere."

— P2, Product Manager

Verification Gap 2

Hallucination detection

"These things are so thirsty to make you happy that they will almost infer intent… as opposed to actually bringing you objective truth."

— P3, Senior UX Designer

these two gaps became the design brief
How might we design a fast, well-integrated verification experience for workers — while maintaining healthy skepticism toward AI-generated meeting artifacts?

/ co-design

/ co-design

"How might WE?" — designing with users

I facilitated a 2-hour online co-design workshop with 4 returning participants. Using a 9-panel storyboard following Jonathan, a senior engineer, we generated ideas together across before, during, and after the meeting. Expanding scope to all phases was a deliberate call: verification failures often originate upstream.

Co-design boards. Participants draw, write, and annotate ideas based on our scenario.

/ concept-testing & the pivot

/ concept-testing & the pivot

Our team tested two concepts with a unanimous result, but I learned something from the unselected version.

Concept A: Verification dashboard — Preferred by all test participants

A full report dashboard that flags items to verify with a confidence scale (high, medium, low)

Concept B: Verification chatbot — not selected

An AI assistant that walks through each uncertain item in conversation and asks follow-ups

Though it was not selected, participants preferred the content design from this option:

"I liked that it said it wasn't sure. It felt more honest than flagging something in red."

— P4, Concept Validation

A business constraint we did not expect: confidence signaling?

"Confidence signaling, while good for users, won't get green lit by leadership. It communicates that our AI is incapable."

— Our mentor from Microsoft

little guy — our mascot — was born!

The animated blob mascot makes the system more friendly and honest.

Asks rather than asserts — hedging language replaces confidence scores

/ the solution

/ the solution

Live Recap Panel

Overview: Live Meeting Recap Panel. Collapsed and expanded states.

Item cards — editable, traceable, flaggable

Decision cards

Flag for review · Source chip · Inline edit

To do cards

Assign to participants in real time · Three states: assigning / assigned / unassigned

End-of-Meeting Co-Review

The host-triggered wrap-up modal is the primary verification moment: two-column layout, all items inline-editable, send immediately or block calendar time for later.

Scrollable recap with summary, decisions, and to-dos

Scrollable recap with Summary, Decisions, and To-Dos. Actions mechanisms are similar to the Live Meeting Recap panel, with the additional Add Item option available to meeting host.

Send recap panel

Send recap panel. Host can select recap participants to send immediately, or choose a time from their calendar to come back and review the recap from their meeting hub.

Post-Meeting Recap Review

Full Recap tab — low-confidence items surfaced

Meeting Recap Review. Confidence signaling + source tracing working together.

/ our impact!

/ our impact!

The workflow reframe

Usability testing confirmed the core hypothesis. Our design delays tool abandonment by inserting a structured, low-cost verification step between low trust and disengagement.

Workflow reframe: traditional four-step flow vs. Teams Up three-step verified-summary flow.

"I like these real-time stuff! I won't have to take notes and it's easy to use!"

— P1, Usability Study, on the Live Recap Panel

"This is really good! I like [that the AI is] looking over and making sure everything is correct! I'd totally use this!"

— P1, Usability Study, on uncertainty flagging and the recap dashboard

/ Beyond the interface — Trustworthy AI Adoption Framework

/ Beyond the interface — Trustworthy AI Adoption Framework

A strategic output for Microsoft and beyond

Trustworthy AI cannot be built through product design alone. We developed an exploratory Plan → Design → Measure framework — delivered to Microsoft XSD as part of the final handoff — for cross-functional teams thinking about AI adoption at every level.

plan

Why adopt, and who's responsible when it goes wrong?

Define purpose, stage rollout so trust is earned rather than assumed, close the accountability gap by naming who owns AI outputs.

Product leaders, TPMs, customer success

design

How to build trust without adding cognitive burden?

Transparency should be lightweight and contextual. Human control over consequential actions is the mechanism through which trust becomes calibrated.

Product designers, researchers, engineers

measure

Is adoption actually working — beyond usage numbers?

High usage can hide falling trust. Verification time that drops may signal over-reliance. Emotional signals are absent from standard product analytics.

PMs, researchers, customer success

/ What I'm Learning

/ What I'm Learning

Some things I've learned along this project

🧐 Research depth changes what you design.

The difference between designing for a brief and designing from evidence is visible in every decision: why the panel is passive, why it doesn't alert, why the host triggers the review.

🤖 Designing for AI requires designing for trust, not just usability.

Standard usability heuristics — efficiency, learnability, error recovery — don't fully account for the trust dynamics in AI systems. A feature can be perfectly usable and still erode appropriate reliance. We're designing for calibration, not just interaction.

🙆🏻‍♀️ Co-design is a great design method, in addition to being a research method!

The co-design workshop generated 20+ ideas in 2 hours that we wouldn't have reached through desk research alone. More importantly, the workshop revealed which problems participants cared about — the before/during/after expansion came directly from what participants wanted to design for.

💡 Speculative work requires more rigor…

Because we're designing for a future AI capability (confidence tiering in Copilot), every design decision needs stronger justification than product work against existing systems. We can't point to a shipped feature — we have to point to research.

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly

/ thank you for stopping by!

Be in touch! I promise I won't bite!

© 2026 by yours truly