Responsible AI measurement

Evaluations

Monitor whether Clarity AI is useful, transparent and safe across real workflows.

Updated 31 Jul · 09:30
Task completion86%+4.1 pts vs prior
Human escalation12.4%−1.8 pts vs prior
Unsupported claims caught97.2%+2.7 pts vs prior
Source coverage94.0%+3.3 pts vs prior
Approval turnaround1.8 days−0.4 days vs prior
User correction rate8.1%+0.6 pts vs prior
Safe-failure success97%+5.0 pts vs prior
Trust rating4.5/5+0.2 vs prior

Quality trends

Monthly quality trend values in percent
MonthTask completionSource coverageSafe-failure success
Feb72%81%84%
Mar75%83%87%
Apr78%86%89%
May77%88%92%
Jun82%91%94%
Jul86%94%97%

Escalations

Most common reasons

Escalation reason counts
ReasonCount
Source conflict38
Policy ambiguity27
Missing market22
Claim unsupported17

Priority insight

Users abandon approval requests when the reason for human review is not explained.

Abandonment is 2.3× higher when the responsible reviewer and expected next step are absent.

Recommendation

Revise the approval pattern to show the risk trigger, responsible reviewer and expected next step.

Repeated failures

Problematic interaction patterns

PatternSignalAffected tasksRecommendationDecision
Approval request · reason omitted19% abandonment142ReviseTest revision
Source conflict · no owner shown31% repeated questions67ReviseOpen
Tool failure · retry only14% task exits51Add manual pathIn progress
Voice confirmation · ambiguous transcript8 accessibility issues23Retire stateReview