Measure & evaluate

Measure how people actually experience your AI.

The System Usability Scale was built for software that behaves the same way every time. AI doesn’t. AI-UX Score is a respondent-backed assessment of how people experience your AI across seven dimensions of trust and usability, the things SUS was never designed to see.

1Trust2Transparency3Agency & Control4Accountability5Fairness6Privacy & Data7Quality & Ease
A sample score across the seven dimensions.
Grounded in UX-research psychometrics Seven trust & usability dimensions First project free · no credit card Export to PDF & CSV
The problem

Standard usability metrics were built for static software. AI is different.

Classic UX research measures how fast someone finishes a task, how many clicks it took, and whether the interface was clear. Those questions still matter, but they assume a system that responds the same way every time. AI is probabilistic, adaptive, and occasionally wrong in ways a user can't predict.

What classic metrics see

  • Task completion and speed
  • Click efficiency
  • Interface clarity
  • Learnability

What AI experience actually turns on

  • Non-deterministic, unexpected outputs
  • Loss of user control
  • Opaque confidence levels
  • Trust erosion after a single bad answer
  • Perceived fairness and data safety

When your product's behaviour is probabilistic, you need a metric built for probability: one that measures perceived trust, control, and accountability, not just whether the button was easy to find.

See it in action

This is what you get back.

Not a spreadsheet of averages, but a living report your team can read at a glance and take straight into planning.

One score · seven dimensions
Overall AI-UX score
4.4/ 7
Moderate / At Risk

Acceptable overall, but real friction is showing. Focus on the weakest dimensions.

24 responses

Trust5.1
Transparency4.2
Agency & Control5.0
Accountability2.9
Fairness4.6
Privacy & Data3.5
Quality & Ease5.8

The shape tells you where trust leaks before you read a word: one headline score, and a seven-axis radar anyone can act on.

Dimension deep-dives
Results by dimension
1Trust5.1 / 7Good
2Transparency4.2 / 7Fair
3Agency & Control5.0 / 7Good
4Accountability2.9 / 7Weak
5Fairness4.6 / 7Fair
6Privacy & Data3.5 / 7Weak
7Quality & Ease5.8 / 7Good

"People don't trust it" becomes "Accountability 2.9, Privacy 3.5." Every dimension gets its own score and grade; open any row to read the questions behind the number.

Track trust over timePro
Overall AI-UX across rounds
1357R1R2R3R4R5R6R7R8R9R104.4

Run it again after every fine-tune or redesign. Ten rounds in, the number that used to be a gut feeling is a line you can defend.

Share & export

Take one clean PDF into a design review or leadership deck, or pull the raw responses into your own analysis. Both come on any plan; the reliability stats that back every number, Cronbach's α and 95% confidence intervals, come with Pro.

PDF reportCSV dataCronbach's α · Pro95% CI · Pro
The payoff

What you actually walk away with.

A score is the start, not the point. Here's what the assessment turns into: a clear diagnosis, and a backlog you can ship.

Payoff 01

Stop guessing which trust problem to fix first.

The seven-axis score shows exactly where confidence is leaking, so your next sprint targets the dimension that's losing users, not the loudest opinion in the room.

Payoff 02

Turn the score into Monday's backlog.

Every weak dimension becomes ranked, concrete design work, grounded in what that dimension measures. A plan to start from, not a verdict to sit with.

  1. 1
    Give a wrong answer somewhere to go.
    Accountability2.9 / 7
  2. 2
    Say what's stored, and whether inputs train the model.
    Privacy & Data3.5 / 7
  3. 3
    Make the AI's confidence legible, and disclose when it's the AI.
    Transparency4.2 / 7
Payoff 03

Prove the redesign worked.

Re-run the assessment after you ship and watch the score move across rounds: the evidence you bring to the next roadmap conversation, not a hunch.

Round 1
Round 10
+1.4

The action plan gives AI-generated suggestions: a starting point to review with your team, not a substitute for talking to users. A Pro feature.

The methodology

The 7 dimensions of AI product experience

AI-UX Score is respondent-backed: real users rate their experience right after using your AI feature. Their answers roll up into seven dimensions, each one a place where trust is either earned or lost.

01 / 07
111TrustFlip the card
Trust

Whether people believe the system is reliable and safe to depend on. It's the difference between acting on a recommendation and quietly ignoring it.

Ask: Do users rely on the output when it actually matters?

222TransparencyFlip the card
Transparency

Whether people can tell what the system is doing, how sure it is, and when they're dealing with AI at all. Status, confidence signals, and clear disclosure live here.

Ask: Can users tell how confident the AI is, and when it's the AI talking?

333Agency & ControlFlip the card
Agency & Control

Whether people can steer, correct, edit, or override the system. Adaptive products fail the moment assistance turns into assumption.

Ask: Can users override the AI when it gets things wrong?

444AccountabilityFlip the card
Accountability

Whether it's clear who is answerable when the AI is wrong, and whether there's a fallback path or a route to recourse.

Ask: When the output is wrong, does the user have somewhere to go?

555FairnessFlip the card
Fairness

Whether the system treats comparable people comparably, and holds output quality steady across different users and inputs.

Ask: Does everyone get output of the same quality?

666Privacy & DataFlip the card
Privacy & Data

Whether people feel their inputs are safe, and whether they understand how their data feeds or trains the model.

Ask: Do users feel safe with what they type in?

777Quality & EaseFlip the card
Quality & Ease

Whether the output is accurate and fluent, and whether getting there costs more effort than the result is worth.

Ask: Is the result good, and worth the effort it took?

How it works

From survey link to actionable redesign in minutes.

1

Launch & collect

Generate a lightweight assessment link, or embed the survey right inside your product flow. Users answer in under five minutes.

2

Benchmark & score

Responses synthesise automatically into your seven-axis radar, with a sub-score for every dimension and a read on how much confidence your sample supports.

3

Identify & act

Get a prioritised gap report: exactly where trust is leaking, and which design changes to make next.

Two clocks, and we keep them honest. Setup takes minutes. Collecting responses takes days, because it depends on real people replying, and that wait is the point.

Plan for about 25 responses in total (a total, not per dimension) to reach a reliable read, and gather them from your users, not your team. A group rating its own work is exactly the self-assessment this instrument exists to replace, which is what makes the number worth defending.

Pricing

Start free. Upgrade when you're tracking.

Free
€0

Your first project, forever.

  • One project, up to 25 responses
  • Score, grade & seven-axis radar
  • Dimension breakdown
  • PDF & CSV export
Start free
Pro
€29 /mo

For teams measuring over time.

  • Unlimited projects, rounds & responses
  • Trends across rounds
  • AI action plan
  • Reliability stats & confidence intervals
  • Shareable hosted reports
Go Pro
Questions

Before you start

How is this different from SUS, and should I use both?

They're companions, not rivals. SUS gives one usability number, validated on conventional software where behaviour is fixed. AI-UX Score is built for probabilistic systems: seven dimensions SUS never covers (trust, transparency, agency, accountability, fairness, data safety, quality). Use SUS for the interface, and AI-UX Score for the AI. Many teams run both. Meet the SUS Calculator →

How many responses do I need?

Around 25 gives a reliable read for most teams, and the tool tells you how much confidence your sample supports. Fewer gives a directional signal; 50–100+ is better for benchmarking over time.

Can I adapt it to my AI domain?

Yes. The seven dimensions are fixed (that keeps scores comparable), but each item has an [AI system] placeholder you set to your context ("this assistant", "the recommendation engine").

Is the free tier really free?

Yes. Your first project is free: one project, up to 25 responses, no card, enough to run a real assessment and get your score and grade across all seven dimensions. Pro lifts the cap to unlimited responses, projects, and rounds.

Ready to see the experience behind your AI?

Set up your first AI-UX survey in under two minutes. No credit card required.

Free first project · up to 25 responses · no credit card