AI-UX Score

Why this exists

After reviewing over 50 AI-powered products and services during my PhD, I found a clear pattern: the systems that succeed long-term don’t just perform well technically — they earn trust, respect user autonomy, and make their logic visible and understandable.

Most teams already track performance metrics like accuracy or latency. Necessary, but not enough to understand the human side of AI. AI-UX Score measures what those metrics can’t: how users feel about interacting with an AI system — whether they trust it, understand it, feel empowered by it, and believe it treats them fairly.

What it measures

Seven dimensions, grounded in behavioral science, human–computer interaction research, and the priorities emphasized in current AI ethics frameworks — the EU AI Act, the OECD AI Principles, and human-centered AI literature.

Trust

Do users believe the system is reliable, predictable, and aligned with their goals?

Trust is the foundation of any successful AI system. It's not enough for users to tolerate AI — they must believe it acts reliably, predictably, and in alignment with their goals and expectations. Trust determines whether users rely on a system in critical moments, or quietly disengage.

Transparency

Can users understand what the system is doing and why?

Users don't trust what they can't understand. Transparency isn't just disclosure — it's making the system's purpose, behavior, and boundaries understandable: how it works, what it's trying to achieve, and what it can and can't do. It reduces friction and builds confidence, and regulation like the EU AI Act is increasingly making it a requirement, not just good practice.

Agency & Control

Do users feel empowered to intervene or adjust the system's behavior?

Effective AI should augment, not override, users. When people feel decisions are made for them without their input, the experience shifts from empowering to alienating. The critical factor is meaningful control — the ability to intervene, override, and set boundaries — especially in adaptive systems where the line between assistance and assumption is thin.

Accountability

Are there clear mechanisms for feedback, oversight, and human responsibility?

Users need to know there are people, not just systems, behind the AI. Accountability means clear feedback loops, escalation paths, and responsibility structures. When users encounter errors or unfair outcomes, they should be able to respond and expect someone to listen — essential for public confidence and regulatory alignment.

Fairness

Is the system perceived as fair, inclusive, and unbiased?

Fairness is not just a model property — it's a user experience. People judge whether an AI is fair based on perceived treatment and inclusion, not only outcomes, and that perception varies across demographics and use cases. Tracking it longitudinally helps catch hidden biases before they become public problems.

Privacy & Data Practices

Do users feel safe and informed about how their data is collected and used?

AI systems increasingly rely on personal, behavioral, or contextual data, but users won't share if they don't feel safe. Privacy isn't just legal compliance — it's a key driver of emotional trust. Users need to know what's collected, how it's used, and whether they can influence that.

Perceived Quality & Ease of Use

Is the system intuitive, helpful, and satisfying in practice?

Even technically advanced AI struggles to gain traction if it's confusing or frustrating. This dimension captures how users evaluate the overall quality of the interaction — not just whether they can use the system, but whether they want to.

What this tool is (and isn’t)

It is

  • Structured insight across 7 dimensions — trust, transparency, control, and fairness, not just one score
  • A user-centered diagnostic — what people feel, not just what they do
  • Actionable and trackable over time or across products
  • Quick to run — most respondents finish in under 5 minutes

It isn’t

  • A universal pass/fail — there's no industry-wide "good" score yet, so comparison over time matters more than the absolute number
  • A substitute for usability testing — pair it with interviews or behavioral data for the full picture
  • Immune to context — domain, user expectations, and system maturity all shape how people respond

Frequently asked questions

How many responses do I need?

There’s no single correct number — it depends on your goal:

  • 20–30 responses: exploratory research, early feedback, pilot studies
  • 50–100 responses: descriptive statistics, benchmarking over time
  • 200+ responses: subgroup comparisons and psychometric validation (factor analysis needs roughly 5–10 responses per item — 115–230+ for this instrument’s 23 items)
Should scores be weighted?

Most of the time, no. Simple, unweighted averages are good enough as long as your sample roughly reflects your real user base. Weighting only helps when your sample is meaningfully unbalanced andyou have reliable reference data about your actual population — otherwise it just adds noise. That’s why AI-UX Score computes plain, unweighted averages by default.

Can I customize the question wording?

Each project has an “AI system label” that’s substituted into every question — that’s the intended customization point. Beyond that, the wording is kept standardized on purpose: it’s been validated as a set, and changing it would affect reliability and comparability, the same reasoning behind not editing standardized instruments like the SUS or UEQ.

What are the limitations?
  • It’s not diagnostic — it tells you how users perceive the experience, not exactly what to fix. Pair it with usability testing or interviews for that.
  • High-performing systems can hit a ceiling effect, where consistently strong scores make small improvements harder to detect.
  • Self-reported answers carry the usual biases — social desirability, the halo effect — mitigated by anonymous, randomized-order responses.
  • Scores are shaped by context: user expectations, domain sensitivity, and how well the system was introduced all move the number independently of the system itself.

Ready to run your own assessment?

Have questions, or want to share how this is helping you design better AI? Reach out on jcerejo.com or on LinkedIn.