Subscore

ReliabilityTesting methodology

Reliability measures how well the chat works without annoying technical problems.

A conversation can feel natural at first, but the experience quickly falls apart when the AI repeats itself, refuses normal messages, takes forever to reply, contradicts earlier details, or sends broken answers.

We also check whether the AI can understand a correction and recover after getting something wrong.

  • 6 evidence groups
  • 6 scored tests

Reliability is organized into 6 evidence groups. Each group contains one scored test: Repetition, Refusals, Reply Speed, Errors, Consistency, and Recovery.

Every test gets a score from 0 to 10.

We multiply each test score by how much it counts. We then add all the points together to calculate the final Reliability score. Reliability makes up of the Chat score.

  • 17%Repetition
  • 17%Refusals
  • 17%Reply Speed
  • 17%Errors
  • 16%Consistency
  • 16%Recovery
Weighted evidence groups (combined 100%)Reliability score of Chat score
View exact calculation

Each test gets a score from 0 to 10. We multiply every score by how much it counts and add the points together.

Reliability makes up of the Chat score. Chat makes up 20% of the overall performance score.

Scored tests and weights

Scored testHow much it countsSee scoring
Repetition17.00%View
Refusals17.00%View
Speed17.00%View
Errors17.00%View
Consistency16.00%View
Recovery16.00%View
Total100%

How the score is calculated

  1. Each test gets a score from 0–10

  2. Test score × how much it counts

    Example: 8 × 17.00% = 1.36 points

  3. We do this for every test

  4. We add all the points together

  5. Final Reliability score

    Counts for 33% of Chat

Example calculation

We multiply each test score by how much it counts. We then add all the points together.

Scored testTest scoreHow much it countsCalculationPoints added
Repetition8.0017.00%8.00 × 17.00%1.36
Refusals8.0017.00%8.00 × 17.00%1.36
Speed8.0017.00%8.00 × 17.00%1.36
Errors9.0017.00%9.00 × 17.00%1.53
Consistency9.2016.00%9.20 × 16.00%1.47
Recovery8.0016.00%8.00 × 16.00%1.28
Final Reliability score100%Add all points8.36/10

Special cases

Not Applicable

If a test does not apply, we remove it and spread its weight across the remaining tests.

Unknown

If we cannot verify a result, the test receives a score of 0.

Manual adjustment

In rare cases, we may adjust a score when the calculated result is clearly unfair. We always record the reason.

Even a good AI girlfriend becomes frustrating when the chat does not work properly.

You may be having a great conversation, and then the AI suddenly repeats the same sentence, forgets what it said five messages earlier, or sends an answer that has nothing to do with the chat.

Slow replies are also annoying. Waiting a few seconds is normal, but regularly waiting 10 or 20 seconds can make the conversation feel dead.

Refusals can be another problem. We are not testing whether the AI allows everything. We send messages that should be allowed under the platform’s own rules and check whether the AI still refuses them for no clear reason.

A reliable chat should respond quickly, avoid repeating itself, keep important details consistent, and recover when you correct a mistake.

We use the same five chats from the Understanding and Realism tests.

Each chat contains 20 AI replies, giving us 100 replies to review.

We check those replies for repeated sentences or ideas, broken or unrelated answers, and contradictions with earlier facts.

We also create one clear misunderstanding in each chat and correct the AI to see whether it can recover.

For the separate refusal test, we send 25 allowed prompts and count how many are refused without a valid reason.

For Reply Speed, we time 25 replies and use the median result. The median is the middle result after the times are placed in order, so one unusually fast or slow reply does not control the score.

Repetition, Refusals, and Errors are displayed per 50 replies or prompts. When the test uses a different sample size, we convert the result to the same public unit so platforms remain easy to compare.

Chat reliability can change depending on your internet connection, device, and the time of day.

A slow reply does not always mean the AI model itself is slow. The platform may be busy, or the connection may be unstable. Timing several replies helps reduce the effect of one unusual result.

Repetition can also be reasonable in some situations. The AI may repeat an important fact when the conversation calls for it. We only count repetition when the same sentence, idea, or reply structure is reused without a clear reason.

A refusal is only counted when the message should be allowed under the platform’s own rules. We do not lower the score because the app refuses content that clearly breaks its policies.

Reliability can also differ between characters. We use five different characters to reduce the effect of one unusually strong or weak chat.

Our results show how the platform performed during the fixed test on the recorded test date.

Evidence groups

Reliability has 6 evidence groups made up of 6 scored tests.

17%

Repetition

1 scored test

Repetition measures how often the AI repeats the same sentence, idea, or type of reply without a good reason.

A repeated phrase once in a long conversation is not a major problem. Repeated answers become frustrating when they make the chat feel stuck or scripted.

Repetition

How often the AI repeats the same sentence, idea, or type of reply without a good reason.

How we test

We review all 100 replies from the five chats. We count replies that repeat the same sentence, the same main idea, the same answer structure, or a previous reply with only a few words changed. We then convert the result to a rate per 50 replies.

What counts
  • Reusing almost the same full reply
  • Repeating the same advice or idea several times
  • Sending the same opening or ending again and again
  • Rewriting an earlier answer without adding anything useful
What does not count
  • Repeating an important detail when it makes sense
  • Referring back to something said earlier
  • Using the same character catchphrase occasionally
  • Similar wording when the answer itself is different
  • Repeating something because the user asked again
Result shown

2 repetition problems per 50 replies

4 repeated replies out of 100

Scoring

Fewer repetition problems means a higher score.

Repetition problems per 50 repliesScore
0 or fewer10/10
19/10
28/10
3–46/10
5–74/10
8–122/10
13+0/10

A result of 2 repetition problems per 50 replies scores 8/10.

Evidence group 2 of 6

17%

Refusals

1 scored test

Refusals measures how often the AI refuses a message that should be allowed.

This does not test whether the platform is completely unfiltered. We only count refusals when the prompt follows the app’s own rules.

Refusals

How often the AI refuses a message that should be allowed.

How we test

We send 25 different prompts that should be allowed under the platform’s current policies. We count how many prompts receive an unnecessary refusal. We then convert the result to a rate per 50 prompts so every platform uses the same scoring unit.

What counts
  • Refusing a harmless request without explaining why
  • Blocking a normal roleplay that follows the rules
  • Refusing a prompt that the platform says is allowed
  • Repeatedly changing the subject instead of responding
What does not count
  • Refusing content that clearly breaks the platform’s rules
  • Warning the user while still answering the allowed part
  • A temporary technical error
  • Asking for clarification when the prompt is unclear
Result shown

2 refusals per 50 prompts

1 unnecessary refusal out of 25 prompts

Scoring

Fewer unnecessary refusals means a higher score.

Refusals per 50 promptsScore
0 or fewer10/10
19/10
28/10
3–46/10
5–74/10
8–122/10
13+0/10

A result of 2 refusals per 50 prompts scores 8/10.

Evidence group 3 of 6

17%

Reply Speed

1 scored test

Reply Speed measures how long the AI normally takes to finish its response.

Fast replies help the conversation feel natural. Long waits can make even a good chat feel slow and awkward.

Reply Speed

How long the AI normally takes to finish its response.

How we test

We time 25 replies. The timer starts when the message is sent and stops when the full AI reply has finished appearing. We use the median reply time.

What counts
  • The full wait from sending the message to the completed reply
  • Normal text replies during the paid test
  • Loading and typing time shown by the platform
What does not count
  • Time spent writing the user’s message
  • Image, voice, or video generation
  • Replies affected by a confirmed internet outage
  • A result recorded before the full response finishes
Result shown

Median reply time: 3.4 seconds

Scoring

Faster replies mean a higher score.

Median reply timeScore
2 or fewer10/10
3–48/10
5–66/10
7–104/10
11–202/10
21+0/10

A median reply time of 3.4 seconds scores 8/10.

Evidence group 4 of 6

17%

Errors

1 scored test

Errors measures how often the AI sends a broken or unusable reply.

This includes answers that are cut off, empty, nonsensical, or completely unrelated to the conversation.

Errors

How often the AI sends a broken or unusable reply.

How we test

We review all 100 replies from the five chats. We count replies that are cut off, empty, broken, nonsensical, or unrelated to the conversation. We then convert the result to errors per 50 replies.

What counts
  • A reply that stops halfway through
  • An empty message
  • Broken formatting that makes the answer unreadable
  • Random or nonsensical text
  • An answer with no clear connection to the chat
What does not count
  • A reply we personally dislike
  • A short answer that still makes sense
  • A factual mistake that does not break the reply
  • A valid refusal, which is tested separately
  • Repetition, which is tested separately
Result shown

1 error per 50 replies

2 errors out of 100 replies

Scoring

Fewer errors means a higher score.

Errors per 50 repliesScore
0 or fewer10/10
19/10
28/10
3–46/10
5–74/10
8–122/10
13+0/10

A result of 1 error per 50 replies scores 9/10.

Evidence group 5 of 6

16%

Consistency

1 scored test

Consistency measures how often the AI contradicts facts already established in the conversation.

The AI may remember a fact correctly at first but later say something that directly conflicts with it.

Consistency

How often the AI contradicts facts already established in the conversation.

How we test

We use the same five personal facts from the Understanding test in each of the five chats. Later in the conversation, we check whether the AI contradicts any of those facts. This creates 25 consistency checks.

What counts
  • Saying the user lives somewhere different
  • Changing the user’s name
  • Referring to the pet by the wrong name
  • Claiming the user has a different favorite food
  • Contradicting the established work schedule
What does not count
  • Forgetting a fact without contradicting it
  • Asking the user to confirm something
  • Correctly updating a fact after the user changes it
  • A vague reply that does not make a clear claim
Result shown

2 contradictions across 25 checks

Contradiction rate: 8%

Scoring

Fewer contradictions means a higher score.

Contradiction rateScore
0%10/10
20%8/10
40%6/10
60%4/10
80%2/10
100%0/10

A contradiction rate of 8% scores 9.2/10.

The exact percentage is used between these points.

Evidence group 6 of 6

16%

Recovery

1 scored test

Recovery measures whether the AI can fix a misunderstanding after you correct it.

Everyone makes mistakes. The important part is whether the AI listens to the correction or continues giving the wrong answer.

Recovery

Whether the AI can fix a misunderstanding after you correct it.

How we test

We create one clear misunderstanding in each of the five chats. We correct the AI immediately afterward. The test passes when the AI understands the correction and responds properly within its next two replies.

What counts
  • Clearly accepting the correction
  • Using the corrected information afterward
  • Fixing the mistake within two replies
  • Continuing the conversation without repeating the same error
What does not count
  • Saying sorry but continuing with the wrong information
  • Ignoring the correction
  • Fixing one detail while repeating the main mistake
  • Only correcting the mistake after being told several times
Result shown

4 of 5 recovery tests passed

Recovery result: 80%

Scoring

A higher recovery rate means a higher score.

Recovery tests passedScore
0 of 50/10
1 of 52/10
2 of 54/10
3 of 56/10
4 of 58/10
5 of 510/10

A result of 4 passed recovery tests scores 8/10.

Back to Chat