Subscore

RealismTesting methodology

Realism measures how natural and human the conversation feels.

An AI girlfriend can remember your name and answer every question correctly, but the chat can still feel robotic. We check whether the replies sound natural, match the character’s personality, handle emotions properly, and help move the conversation forward.

  • 6 evidence groups
  • 6 scored tests

Realism is organized into 6 evidence groups. Each group contains one scored test: Naturalness, Personality, Roleplay, Initiative, Emotion, and Style.

Every test gets a score from 0 to 10.

We multiply each test score by how much it counts. We then add all the points together to calculate the final Realism score. Realism makes up of the Chat score.

  • 17%Naturalness
  • 17%Personality
  • 17%Roleplay
  • 17%Initiative
  • 16%Emotion
  • 16%Style
Weighted evidence groups (combined 100%)Realism score of Chat score
View exact calculation

Each test gets a score from 0 to 10. We multiply every score by how much it counts and add the points together.

Realism makes up of the Chat score. Chat makes up 20% of the overall performance score.

Scored tests and weights

Scored testHow much it countsSee scoring
Naturalness17.00%View
Personality17.00%View
Roleplay17.00%View
Initiative17.00%View
Emotion16.00%View
Style16.00%View
Total100%

How the score is calculated

  1. Each test gets a score from 0–10

  2. Test score × how much it counts

    Example: 8.40 × 17.00% = 1.43 points

  3. We do this for every test

  4. We add all the points together

  5. Final Realism score

    Counts for 33% of Chat

Example calculation

We multiply each test score by how much it counts. We then add all the points together.

Scored testTest scoreHow much it countsCalculationPoints added
Naturalness8.4017.00%8.40 × 17.00%1.43
Personality8.0017.00%8.00 × 17.00%1.36
Roleplay8.4017.00%8.40 × 17.00%1.43
Initiative7.8017.00%7.80 × 17.00%1.33
Emotion8.4016.00%8.40 × 16.00%1.34
Style8.2016.00%8.20 × 16.00%1.31
Final Realism score100%Add all points8.20/10

Special cases

Not Applicable

If a test does not apply, we remove it and spread its weight across the remaining tests.

Unknown

If we cannot verify a result, the test receives a score of 0.

Manual adjustment

In rare cases, we may adjust a score when the calculated result is clearly unfair. We always record the reason.

Chat is the main reason most people use an AI girlfriend app.

The conversation should not feel like you are talking to a customer support bot wearing a cute profile picture.

A realistic AI girlfriend should have her own personality, react properly to your mood, and help keep the conversation going. She should not send the same safe reply every time or make you do all the work.

Roleplay is also a big part of these apps. A good character should add details, respond to your actions, and keep the story moving instead of giving short and lifeless answers.

That is why we do not only check whether the AI understands you. We also check whether talking to it actually feels natural and enjoyable.

We use the same five chats from the Understanding test.

Each chat uses a different character and includes 20 AI replies. This gives us 100 replies to review.

For every chat, we record how many of the 20 replies sound natural, whether the character keeps her personality, how many of the five roleplay checks pass, how often the character takes useful initiative, how well it handles five emotional moments, and how many replies match the selected communication style.

Using the same chats keeps the tests consistent and lets us judge several parts of the conversation without starting with a completely different set of characters.

Realism is partly based on judgement. Different users may prefer different reply lengths, writing styles, or levels of detail.

We reduce this problem by using the same checks for every app.

Chat quality can also change between characters. One character may feel very natural while another feels robotic, even when they use the same platform.

We test five different characters to reduce the effect of one unusually strong or weak character.

Realism also changes depending on the conversation. An AI may perform well during casual chat but struggle with emotional topics or longer roleplays.

Our results show how the platform performed during the fixed test on the recorded test date.

Evidence groups

Realism has 6 evidence groups made up of 6 scored tests.

17%

Naturalness

1 scored test

Naturalness measures how many replies sound like something a real person might actually send.

A reply does not need to be perfect. It should simply feel natural for the conversation instead of sounding robotic, copied, or strangely formal.

Naturalness

How many replies sound like something a real person might actually send.

How we test

We review all 100 replies from the five chats. Every reply is checked for natural wording, a suitable reply length, logical flow, and no robotic or copy-paste language. A reply passes when it meets at least three of these four checks.

What counts
  • Wording that sounds natural
  • A reply length that fits the message
  • A response that flows from the previous message
  • Casual language when it suits the character
  • Longer replies when the situation needs more detail
What does not count
  • Robotic or overly formal wording
  • Replies that feel copied from a template
  • Very long essays in response to a simple message
  • One-line answers when more detail is clearly needed
  • Replies that suddenly change the subject
Result shown

84 of 100 replies passed

Naturalness result: 84%

Scoring

A higher percentage means a higher score.

Natural repliesScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 84% scores 8.4/10.

The exact percentage becomes the score.

Evidence group 2 of 6

17%

Personality

1 scored test

Personality measures whether the character keeps the traits she is supposed to have.

A bratty goth character should not suddenly turn into a polite life coach after ten messages.

Personality

Whether the character keeps the traits she is supposed to have.

How we test

Each of the five tested characters has three clear personality traits. We review all 20 replies in each chat. A chat passes when the character keeps at least two of the three traits throughout the conversation.

What counts
  • The character’s wording matches her personality
  • Her reactions make sense for the selected traits
  • The personality stays clear across the full chat
  • Small changes in mood that still fit the character
What does not count
  • Mentioning a trait once and then ignoring it
  • Becoming a completely different person halfway through
  • Generic replies that could come from any character
  • Breaking personality whenever the topic changes
Result shown

4 of 5 chats kept the selected personality

Personality result: 80%

Scoring

A higher pass rate means a higher score.

Chats that passScore
0 of 50/10
1 of 52/10
2 of 54/10
3 of 56/10
4 of 58/10
5 of 510/10

A result of 4 passed chats scores 8/10.

Evidence group 3 of 6

17%

Roleplay

1 scored test

Roleplay measures how enjoyable and well-written the roleplay feels.

Understanding tests whether the AI remembers the basic scenario. Realism checks whether it actually does something interesting with it.

Roleplay

How enjoyable and well-written the roleplay feels.

How we test

We review the roleplay inside all five chats. Each chat receives one point for every check it passes: stays in character, adds useful details, responds to the user’s actions, keeps the story consistent, and moves the scenario forward. This creates 25 roleplay checks in total.

What counts
  • Describing useful actions or details
  • Reacting to what the user does
  • Keeping the story consistent
  • Adding something new to the scene
  • Giving the user something useful to respond to
What does not count
  • Repeating the user’s message
  • Giving very short replies that add nothing
  • Ignoring actions inside the roleplay
  • Suddenly changing the location or story
  • Waiting for the user to control every part of the scene
Result shown

21 of 25 roleplay checks passed

Roleplay result: 84%

Scoring

A higher percentage means a higher score.

Roleplay checks passedScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 84% scores 8.4/10.

The exact percentage becomes the score.

Evidence group 4 of 6

17%

Initiative

1 scored test

Initiative measures how often the character helps move the conversation forward.

A good AI girlfriend should not make you ask every question and introduce every new topic yourself.

Initiative

How often the character helps move the conversation forward.

How we test

We use 10 open-ended messages in each of the five chats. This creates 50 chances for the character to take initiative. A reply passes when it asks a relevant question, adds a useful new detail, or suggests a logical next action.

What counts
  • Asking a question that fits the topic
  • Introducing a useful new detail
  • Suggesting something to do next
  • Moving the roleplay forward
  • Keeping the conversation alive naturally
What does not count
  • Asking a random question to fill space
  • Repeating the user’s last message
  • Adding details that do not fit the conversation
  • Changing the subject for no reason
  • Ending every reply without giving the user anything to respond to
Result shown

39 of 50 replies showed useful initiative

Initiative result: 78%

Scoring

A higher percentage means a higher score.

Replies showing initiativeScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 78% scores 7.8/10.

The exact percentage becomes the score.

Evidence group 5 of 6

16%

Emotion

1 scored test

Emotion measures how well the AI responds to the user’s mood.

A realistic character should react differently when you are happy, sad, angry, nervous, or romantic.

Emotion

How well the AI responds to the user’s mood.

How we test

We use five emotional situations in each of the five chats: happy, sad, angry, nervous, and romantic. This creates 25 emotional-response tests. A reply passes when the character notices the mood and responds in a suitable way.

What counts
  • Matching the user’s emotional tone
  • Showing support when the user is upset
  • Reacting naturally to good news
  • Handling anger without ignoring it
  • Responding properly to romantic messages
What does not count
  • Giving the same reply to every emotion
  • Ignoring an obvious emotional cue
  • Making a joke during a serious moment without a good reason
  • Becoming romantic when the user is upset
  • Sending a generic reply that could fit any mood
Result shown

21 of 25 emotional moments were handled well

Emotion result: 84%

Scoring

A higher percentage means a higher score.

Emotional moments handled wellScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 84% scores 8.4/10.

The exact percentage becomes the score.

Evidence group 6 of 6

16%

Style

1 scored test

Style measures whether the replies match the communication style selected for the character.

For example, a short and playful style should not suddenly turn into long and formal essays.

Style

Whether the replies match the communication style selected for the character.

How we test

We select one communication style for each of the five chats. We review all 20 replies in every conversation. A reply passes when it clearly matches the selected style. This creates 100 style checks in total.

What counts
  • Reply length that matches the selected style
  • Wording that fits the chosen tone
  • The style stays clear throughout the chat
  • Natural changes that still fit the same overall style
What does not count
  • The selected style only appearing in the first reply
  • Every character sounding exactly the same
  • Sudden changes from casual to formal
  • Long replies when a short style was selected
  • Generic replies with no clear communication style
Result shown

82 of 100 replies matched the selected style

Style result: 82%

Scoring

A higher percentage means a higher score.

Replies matching the styleScore
20%2/10
40%4/10
60%6/10
80%8/10
100%10/10

A result of 82% scores 8.2/10.

The exact percentage becomes the score.

Back to Chat