Subscore

UnderstandingTesting methodology

Understanding measures how well the AI follows what you say during a conversation.

A reply can sound good at first but still be frustrating if the AI forgets your name, ignores your question, misses something you said earlier, or completely misunderstands the roleplay.

We test whether the AI remembers details, gives direct answers, follows the conversation, listens to your instructions, and understands the scenario.

  • 5 evidence groups
  • 5 scored tests

Understanding is organized into 5 evidence groups. Each group contains one scored test: Memory, Relevance, Context, Instructions, and Roleplay Accuracy. All five tests count equally toward the final Understanding score.

Every scored test gets a score from 0 to 10.

All five tests count equally. We combine their scores to calculate the final Understanding score. Understanding makes up of the Chat score.

  • 20%Memory
  • 20%Relevance
  • 20%Context
  • 20%Instructions
  • 20%Roleplay Accuracy
Weighted evidence groups (combined 100%)Understanding score of Chat score
View exact calculation

Each test gets a score from 0 to 10. We multiply every score by 20% and add the points together.

Understanding makes up of the Chat score. Chat makes up 20% of the overall performance score.

Scored tests and weights

Scored testHow much it countsSee scoring
Memory20.00%View
Relevance20.00%View
Context20.00%View
Instructions20.00%View
Roleplay Accuracy20.00%View
Total100%

How the score is calculated

  1. Each test gets a score from 0–10

  2. Test score × how much it counts

    Example: 8.40 × 20.00% = 1.68 points

  3. We do this for every test

  4. We add all the points together

  5. Final Understanding score

    Counts for 34% of Chat

Example calculation

We multiply each test score by how much it counts. We then add all the points together.

Scored testTest scoreHow much it countsCalculationPoints added
Memory8.4020.00%8.40 × 20.00%1.68
Relevance8.8020.00%8.80 × 20.00%1.76
Context8.0020.00%8.00 × 20.00%1.60
Instructions8.6720.00%8.67 × 20.00%1.73
Roleplay Accuracy8.8020.00%8.80 × 20.00%1.76
Final Understanding score100%Add all points8.53/10

Special cases

Not Applicable

If a test does not apply, we remove it and spread its weight across the remaining tests.

Unknown

If we cannot verify a result, the test receives a score of 0.

Manual adjustment

In rare cases, we may adjust a score when the calculated result is clearly unfair. We always record the reason.

Talking to an AI girlfriend becomes annoying very quickly when she does not understand you.

You might tell her your name, where you live, and what you like. A few messages later, she may forget everything and ask the same questions again.

The same problem happens with instructions. You may ask for short replies, tell her not to use emojis, or set up a specific roleplay. A weak AI ignores those rules and does whatever it wants.

A good chat should do more than create nice-sounding sentences. It should understand what you are asking, remember what has already happened, and respond in a way that makes sense for the conversation.

We open five new chats with five different characters.

We use the same script in every chat so that each platform receives the same test.

In every conversation, we give the AI five facts about ourselves, ask five direct questions, set three simple rules, build a short conversation that requires earlier context, and start the same roleplay with five checks.

This gives us 25 memory checks, 25 direct-question checks, 5 context results, 15 instruction checks, and 25 roleplay checks.

We enter one row for each of the five chats in our testing worksheet.

The five facts are: my name is Herman, I live in Bangkok, my favorite food is pizza, I have a dog named Milo, and I work at night.

The three rules are: call me Herman, keep replies under three sentences, and do not use emojis.

The five direct questions cover a rainy date, favorite movie, cheering up after a bad day, dinner, and vacation.

The roleplay starts at a quiet hotel bar. The character is confident but slightly nervous, stays in character, and describes actions in italics.

Chat quality can change between characters. One character may understand you very well while another performs much worse on the same platform.

We use five different characters to reduce this problem, but we cannot test every character in the library.

This test measures understanding inside the tested conversations. It does not prove that the AI will remember everything after several days or weeks.

Manual memory tools are also tested separately under Chat Features. Understanding focuses on whether the AI naturally remembers and uses information during the conversation.

AI models change regularly. Our results show how the chat performed on the recorded test date.

Evidence groups

Understanding has 5 evidence groups made up of 5 scored tests.

20%

Memory

1 scored test

Memory measures how many personal facts the AI remembers later in the conversation.

A chat feels much more personal when the AI remembers basic things about you instead of asking the same questions again.

Memory

How many personal facts the AI remembers later in the conversation.

How we test

We give the AI five facts in each of the five chats. Later in the conversation, we ask questions to see whether it still remembers them. The five facts cover name, location, favorite food, pet, and work schedule.

What counts
  • The correct fact
  • A clearly correct answer using slightly different wording
  • A natural reference to the fact without being reminded
  • An answer that keeps the important meaning correct
What does not count
  • A wrong answer
  • A vague guess
  • Repeating the fact immediately after we give it
  • Remembering one part while changing the important detail
  • Information saved manually through a memory tool
Result shown

21 of 25 facts remembered

Memory result: 84%

Scoring

A higher percentage means a higher score.

Facts rememberedScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 84% scores 8.4/10.

The exact percentage becomes the score.

Evidence group 2 of 5

20%

Relevance

1 scored test

Relevance measures how often the AI gives a clear answer to the question you asked.

A reply can sound natural but still be useless when it avoids the question or changes the subject.

Relevance

How often the AI gives a clear answer to the question you asked.

How we test

We ask five direct questions in each of the five chats. A reply passes when it clearly answers the question and stays on topic.

What counts
  • A clear answer to the question
  • An answer that adds useful details
  • A short or long answer when the main question is still answered
  • A relevant follow-up question after answering
What does not count
  • Changing the subject
  • Giving a vague answer that avoids the question
  • Only asking a question back
  • Sending an unrelated response
  • Ignoring an important part of the question
Result shown

22 of 25 questions answered directly

Relevance result: 88%

Scoring

A higher percentage means a higher score.

Direct answersScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 88% scores 8.8/10.

The exact percentage becomes the score.

Evidence group 3 of 5

20%

Context

1 scored test

Context measures whether the AI can use information from earlier in the same conversation.

A weak AI only reacts to the most recent message. A stronger AI understands how the newest message connects to what happened before.

Context

Whether the AI can use information from earlier in the same conversation.

How we test

We create a short multi-message exchange in each of the five chats. Later, we send a message that only makes sense when the AI remembers and understands the earlier parts of the conversation. Each chat receives one Context result.

What counts
  • Correctly using something said earlier
  • Remembering the setting or topic
  • Understanding how several messages connect
  • Responding based on the full conversation
What does not count
  • Only reacting to the latest message
  • Giving a generic reply that could fit any conversation
  • Guessing correctly without using earlier details
  • Contradicting something established earlier
Result shown

4 of 5 chats used earlier context correctly

Context result: 80%

Scoring

A higher percentage means a higher score.

Context tests passedScore
0 of 50/10
1 of 52/10
2 of 54/10
3 of 56/10
4 of 58/10
5 of 510/10

A result of 4 passed chats scores 8/10.

Evidence group 4 of 5

20%

Instructions

1 scored test

Instructions measures how well the AI follows simple rules you give it.

This matters when you want a specific reply style, format, or behavior during the chat.

Instructions

How well the AI follows simple rules you give it.

How we test

We give the same three rules in each of the five chats: call me Herman, keep replies under three sentences, and do not use emojis. We check each rule separately. This creates 15 instruction checks in total.

What counts
  • Calling the user Herman when appropriate
  • Keeping replies within the requested length
  • Avoiding emojis
  • Following the rules throughout the test
What does not count
  • Following a rule once and then repeatedly breaking it
  • Ignoring one rule because the other two were followed
  • Almost following the request
  • A lucky reply that follows the rule without showing consistency
Result shown

13 of 15 instructions followed

Instructions result: 86.7%

Scoring

A higher percentage means a higher score.

Instructions followedScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 86.7% scores 8.67/10.

The exact percentage becomes the score.

Evidence group 5 of 5

20%

Roleplay Accuracy

1 scored test

Roleplay Accuracy measures whether the AI understands the scenario and keeps the important details correct.

A roleplay quickly falls apart when the character forgets the location, changes roles, or responds as if the scene never happened.

Roleplay Accuracy

Whether the AI understands the scenario and keeps the important details correct.

How we test

We use the same hotel-bar roleplay in all five chats. Each conversation receives one point for every check it passes: starts the scenario correctly, stays in character, remembers the setting, responds properly to actions, and does not contradict or break the scene. This creates 25 roleplay checks in total.

What counts
  • Correctly starting in the hotel bar
  • Keeping the assigned personality and role
  • Remembering who the user is
  • Responding to actions inside the scene
  • Keeping the situation consistent
What does not count
  • Moving to a new location without a reason
  • Changing the user’s or character’s role
  • Ignoring actions written in the roleplay
  • Speaking out of character
  • Breaking the scene or contradicting earlier events
Result shown

22 of 25 roleplay checks passed

Roleplay Accuracy result: 88%

Scoring

A higher percentage means a higher score.

Roleplay checks passedScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 88% scores 8.8/10.

The exact percentage becomes the score.

Back to Chat