UX Metrics

Beyond Accuracy: How to Benchmark AI User Experience

Joana CerejoAlso on Anticipatory Design Lab
Anticipatory DesignPredictive UXBusiness StrategyAI

TL;DR

  • Technical accuracy keeps rising while public trust in AI falls; benchmarks that only measure performance miss what actually drives adoption.
  • A 23-item, respondent-backed framework scores AI experience across trust, transparency, and agency, the human factors traditional KPIs ignore.
  • Treat trust as a measurable design outcome, not a lucky side effect.

A 23-item evaluation framework for Product Managers and Designers to measure trust, transparency, and agency in human-centered AI systems.

As artificial intelligence (AI) continues to reshape industries, organizations face a growing paradox: technical capabilities keep improving, yet public trust in AI is declining. A Forbes research found that user trust in AI fell from 61% to 54% in just one year.

This is more than a perception problem; it’s a strategic risk.

Traditional performance metrics — like accuracy, latency, or throughput — are essential for understanding how well an AI system performs technically, but they tell only part of the story. A chatbot might respond with 99% accuracy, but if users find its tone cold, its explanations unclear, or its decisions opaque, users may quickly lose trust. These human factors — trust, clarity, perceived fairness, and emotional resonance — are what ultimately determine whether people embrace or abandon an AI system.

User adoption isn’t won by technical performance alone; it’s earned through a consistently positive, trustworthy experience. Traditional KPIs miss what matters most to users:

Do they trust the AI’s actions? Do they feel empowered or sidelined by automation? Is the AI fair and understandable, or does it feel like a black box?

Filling the Human-Centered AI Measurement Gap

Most KPIs focus on what’s easy to measure: system performance, error rates, or model accuracy. But after reviewing over 50 AI-powered products during my PhD research, I found a common pattern: the most successful systems balance three core dimensions — trust, agency, and transparency.

These factors remain invisible to traditional technical metrics.

Many organizations already make efforts to measure their AI’s perceived quality, but what’s missing in the market is a common, standardized, and reliable way to measure it systematically — enabling the creation of a global benchmark.

That’s why I developed the AI-UX Self-Assessment Questionnaire.

https://payhip.com/b/nWMmL Benchmark AI Experiences to Drive Meaningful Adoption This questionnaire is more than a health check. It’s a strategic benchmark for teams building or scaling AI solutions. It helps you:

Understand why technical performance alone isn’t driving adoption Benchmark user perceptions across trust, transparency, fairness, and agency — core to responsible, scalable AI Prioritize investments like explainability or override controls based on what truly impacts user trust Meet governance, compliance, and ethical review requirements as human-centered AI becomes a regulatory and reputational imperative The value of the questionnaire isn’t just in diagnosis — it’s in providing actionable insight for design, product development, and oversight.

The Seven Dimensions of AI-UX

Each dimension is rooted in practical UX research and reflects the priorities found in leading AI ethics and design frameworks (e.g., EU AI Act, OECD AI Principles, and recent HCAI literature).

1. Trust

Trust is the foundation of any successful AI system. It’s not enough for users to tolerate AI — they must believe it acts reliably, predictably, and in alignment with their goals and expectations.

As user trust in AI continues to decline globally, organizations that take the time to diagnose and address trust barriers will have a significant long-term advantage. Trust determines whether users rely on a system in critical moments — or quietly disengage.

Question example: I trust the recommendations or decisions made by the [AI system].

2. Transparency

Users don’t trust what they can’t understand. Transparency isn’t just a matter of disclosure — it’s about making the system’s purpose, behavior, and boundaries understandable. This includes explaining how the AI works, what it’s trying to achieve, and what it can and can’t do.

Transparency reduces cognitive friction, builds confidence, and increases user readiness to rely on AI support. In fact, regulatory guidelines — like the EU AI Act — are increasingly making transparency not just good practice, but a legal requirement.

Question example: It’s clear when I’m interacting with the [AI system].

3. Agency & Control

Effective AI should augment — not override — users. When people feel that decisions are being made for them without their input, the experience quickly shifts from empowering to alienating.

The critical factor is meaningful control — the ability to intervene, override, and set boundaries. This becomes especially important in adaptive or anticipatory systems, where the line between assistance and assumption is thin. If the system acts without clarity or consent, users may experience it as coercive — even if unintentionally so. This perceived loss of agency is a well-documented reason for user rejection, particularly when automation begins to take over action rather than support it.

To sustain trust, autonomy must be shaped by the user’s perspective — not by what the AI technology is capable of doing. [1] [2] [3]

Question example: I feel I have meaningful control over the [AI system’s] actions.

4. Accountability

Users need to know that there are people — not just systems — behind the AI. Accountability is about creating clear feedback loops, escalation paths, and responsibility structures.

When users encounter errors, confusion, or unfair outcomes, they should be able to respond — and expect someone to listen. This not only supports trust but also enables the continuous improvement of AI through real-world feedback. Knowing that someone is ultimately accountable is essential for public confidence and regulatory alignment.

Question example: There is a clear way to provide feedback on the AI’s performance and decisions.

5. Fairness

Fairness is not just a model property — it’s a user experience. People assess whether an AI is fair not only based on outcomes, but on perceived treatment, inclusion, and bias.

This perception varies across demographics, roles, and use cases. Organizations that track fairness longitudinally and across user groups are better positioned to detect hidden biases before they become public scandals or regulatory violations. Social legitimacy depends on people believing that the AI respects their context and treats them equitably.

Question example: The [AI system] treats all users fairly, regardless of their background, group, or role.

6. Privacy & Data Practices

AI systems increasingly rely on personal, behavioral, or contextual data. But users won’t share if they don’t feel safe.

Privacy isn’t just about legal compliance — it’s a key driver of emotional trust. Users need to know what data is being collected, how it’s being used, and whether they can influence that process. Especially in AI systems that adapt or personalize, clarity around data boundaries becomes critical for sustained engagement.

Question example: The [AI system] clearly informs me about what (personal/work) data is collected and why.

7. Perceived Quality & Ease of Use

Even technically advanced AI systems will struggle to gain traction if they’re confusing, frustrating, or fail to meet user expectations in real-world settings. But ease of use alone isn’t enough.

This dimension captures how users evaluate the overall quality of their interaction with the system — how intuitive, helpful, responsive, and satisfying it feels. It reflects not just whether users can use the system, but whether they want to. High perceived quality often correlates with future engagement, especially when users face edge cases, ambiguous feedback, or unexpected system behavior.

Question example: I am satisfied with my experience using the [AI system].

screenshot of study results from excel calculator. When should you run an AI-UX Questionnaire Study? You should run an AI-UX Questionnaire study at key moments when you’re looking to evaluate or improve the user experience of an AI-powered system. This includes:

Before or after launching a new AI feature, to gather feedback on user perceptions and readiness. During usability testing, to supplement qualitative insights with structured perception data. At regular intervals, to monitor how trust, control, and satisfaction evolve over time. After major updates, especially if they affect automation, personalization, or decision-making logic. When adoption is low or user feedback is unclear, to identify hidden friction points or trust barriers. Can other factors affect the AI-UX score? Yes, several external and contextual factors can influence the AI-UX Questionnaire score beyond the system’s actual design or performance. These include:

User Expectations and Prior Experience: Users with past negative experiences with AI — or very high expectations — may rate a system more critically, regardless of actual performance. System Maturity: Early-stage or experimental AI systems may receive lower scores due to bugs, limited capabilities, or lack of refinement. Use Case and Domain Sensitivity: AI in high-stakes domains (e.g., healthcare, finance) tends to be judged more strictly than in casual or low-risk applications (e.g., entertainment, shopping). Cultural and Demographic Differences: Perceptions of trust, fairness, and privacy vary across cultural and social contexts, influencing how users respond. Communication and Onboarding: If the system is poorly introduced or lacks clear explanations, users may feel confused or skeptical — even if the AI performs well. Device and Interface Constraints: A good AI experience on one platform (e.g., desktop) may not translate to another (e.g., mobile), skewing scores based on access channel. Recent Incidents or Media Coverage: Public perception is easily influenced by media coverage or scandals. If you conduct a study where such events have occurred recently, they may interactively influence your participants’ responses, even if your AI system is committed to compliance and ethics. Curious to learn more? I have prepared a comprehensive Notion toolkit where you can find the full AI-UX Self-Assessment Questionnaire, which includes:

The full 23-item AI-UX Questionnaire covering all seven dimensions A detailed scoring guide and interpretation model Multiple questionnaire templates (PDF, Google Forms, Tally) An FAQ to address common questionnaire and quantitative research questions An Excel calculator to automate scoring and analysis Explore the toolkit [available here], or if you just want the questionnaire [find it here].

https://payhip.com/b/xONpL

Why This Matters

Technical excellence doesn’t guarantee adoption. Ethical compliance doesn’t guarantee trust. Governance doesn’t guarantee good design.

The future of AI depends on delivering experiences that align with user needs, expectations, and values. The AI-UX Self-Assessment Questionnaire gives teams a practical way to measure and improve those experiences — with people at the center.

Let’s build AI that people can trust

If you’d like to try the questionnaire in your team or product, reach out. I’d love to hear how it works for you — and how we can keep improving it together.

If you want a deep dive into these 7 dimensions, my book “The Anticipatory Design Playbook” will be released later this year.

Share
Try it in the toolkit

AI-UX Score

Find out whether people actually trust your AI: a respondent-backed score across 7 dimensions of experience.