Measure & evaluate

Measure how people actually experience your AI.

The System Usability Scale was built for software that behaves the same way every time. AI doesn’t. AI-UX Score is a respondent-backed assessment of how people experience your AI across seven dimensions of trust and usability, the things SUS was never designed to see.

1Trust2Transparency3Agency & Control4Accountability5Fairness6Privacy & Data7Quality & Ease
A sample score across the seven dimensions.
The problem

Standard usability metrics were built for static software. AI is different.

Classic UX research measures how fast someone finishes a task, how many clicks it took, and whether the interface was clear. Those questions still matter, but they assume a system that responds the same way every time. AI is probabilistic, adaptive, and occasionally wrong in ways a user can’t predict.

What classic metrics see

  • Task completion and speed
  • Click efficiency
  • Interface clarity
  • Learnability

What AI experience actually turns on

  • Non-deterministic, unexpected outputs
  • Loss of user control
  • Opaque confidence levels
  • Trust erosion after a single bad answer
  • Perceived fairness and data safety

When your product’s behaviour is probabilistic, you need a metric built for probability: one that measures perceived trust, control, and accountability, not just whether the button was easy to find.

The methodology

The 7 dimensions of AI product experience

AI-UX Score is respondent-backed: real users rate their experience right after using your AI feature. Their answers roll up into seven dimensions, each one a place where trust is either earned or lost.

  • Trust

    Whether people believe the system is reliable and safe to depend on. It's the difference between acting on a recommendation and quietly ignoring it.

    Ask: Do users rely on the output when it actually matters?

  • Transparency

    Whether people can tell what the system is doing, how sure it is, and when they're dealing with AI at all. Status, confidence signals, and clear disclosure live here.

    Ask: Can users tell how confident the AI is, and when it's the AI talking?

  • Agency & Control

    Whether people can steer, correct, edit, or override the system. Adaptive products fail the moment assistance turns into assumption.

    Ask: Can users override the AI when it gets things wrong?

  • Accountability

    Whether it's clear who is answerable when the AI is wrong, and whether there's a fallback path or a route to recourse.

    Ask: When the output is wrong, does the user have somewhere to go?

  • Fairness

    Whether the system treats comparable people comparably, and holds output quality steady across different users and inputs.

    Ask: Does everyone get output of the same quality?

  • Privacy & Data

    Whether people feel their inputs are safe, and whether they understand how their data feeds or trains the model.

    Ask: Do users feel safe with what they type in?

  • Quality & Ease

    Whether the output is accurate and fluent, and whether getting there costs more effort than the result is worth.

    Ask: Is the result good, and worth the effort it took?

How it works

From survey link to actionable redesign in minutes.

1

Launch & collect

Generate a lightweight assessment link, or embed the survey right inside your product flow. Users answer in under five minutes.

2

Benchmark & score

Responses synthesise automatically into your seven-axis radar, with a sub-score for every dimension and a read on how much confidence your sample supports.

3

Identify & act

Get a prioritised gap report: exactly where trust is leaking, and which design changes to make next.

The output

What your team gets

  • Aggregate score & grade

    One headline number, graded A+ to F, so anyone can read the state of trust at a glance.

  • Seven-axis radar

    Your UX debt made visible: the shape of the chart tells you where to look before you read a word.

  • Dimensional deep dives

    Granular scores per dimension, so “people don't trust it” becomes “agency scores 4.1, transparency 6.6.”

  • Actionable redesign backlog

    Low scores translated into concrete UI and design tasks: a backlog to work through, not a verdict to sit with.

  • Shareable PDF & CSV export

    Export for design reviews, stakeholders, and leadership, or pull the raw data into your own analysis.

Who it’s for

Built for every stage of the AI lifecycle

  • Product designers

    Validate prototype concepts and steerability controls before engineering handoff, while changes are still cheap.

  • AI product managers

    Set a baseline trust score, then measure again after a fine-tune or a major update. Know whether the change earned trust or eroded it.

  • UX researchers

    Swap generic survey tools for a validated, AI-native framework, with the psychometrics to defend the numbers.

Stop guessing if users trust your AI.

Measuring trust isn’t a one-off audit, and it isn’t a step in a pipeline. It’s one of four questions every AI team keeps asking, in no fixed order. What you anticipate, who you map, and what you must comply with all shape what “trust” even means here. AI-UX Score is where you measure it, and each question makes the others sharper.

Questions

Before you start

  • How does the AI-UX Score differ from the System Usability Scale (SUS)?

    SUS gives you one number for general usability, validated on conventional software where behaviour is fixed. AI-UX Score is purpose-built for probabilistic systems: it measures seven dimensions SUS never covers (trust, transparency, agency, accountability, fairness, data safety, and quality) because with AI, “usable” and “trusted” are not the same thing. Use SUS for the interface; use AI-UX Score for the AI.

  • How many user responses do I need to get a statistically useful score?

    Around 25 gives a reliable read for most teams, and the tool tells you how much confidence your sample supports. Fewer still gives a directional signal for early work; 50–100+ is better for benchmarking over time or across products; and 150–200+ if you plan to run subgroup comparisons or reliability analysis.

  • Can I customise the questions for my specific AI domain?

    Yes. The seven dimensions are fixed (that's what keeps scores comparable) but the items are written with an [AI system] placeholder you adapt to your context (“this assistant”, “the recommendation engine”). Keep each item's intent intact and your results stay valid, whether you're assessing generative text or predictive analytics.

  • Is this suitable for internal B2B tools as well as customer-facing B2C AI products?

    Both. Trust, control, and accountability matter just as much for an internal ops copilot as a consumer app, arguably more, since the people relying on it can't simply churn. The framework is domain-agnostic; you set the context.

  • Is the free assessment tier really free?

    Yes. Your first project is free, with unlimited responses and no card required: enough to run a real assessment and get your score and grade across all seven dimensions. Pro adds unlimited projects, trends over time, reliability stats, and shareable reports.

Want the full methodology: sample sizing, weighting, and the maths? Read the complete FAQ →

Ready to measure the experience behind your AI?

Set up your first AI-UX survey in under two minutes. No credit card required.

Unlimited responses on the free tier · Export results anytime