Subscore

InteractionTesting methodology

Interaction measures how much the AI can do beyond sending one normal reply at a time.

We check whether you can make voice calls, use different chat modes, create group chats, receive multiple messages, and get messages from the character without writing first.

These features can make the app feel much more like a real relationship and less like a basic chatbot.

  • 6 evidence groups
  • 6 scored tests

Interaction is organized into 6 evidence groups. Each group contains one scored test: Voice Calls, Chat Modes, Mode Types, Group Chat, Double Texting, and Proactive Messages.

Every test gets a score from 0 to 10.

We multiply each test score by how much it counts. We then add all the points together to calculate the final Interaction score.

Voice Calls has the highest weight. Group Chat has the lowest weight. Interaction makes up of the Chat Features score.

  • 27%Voice Calls
  • 22%Chat Modes
  • 17%Mode Types
  • 5%Group Chat
  • 15%Double Texting
  • 14%Proactive Messages
Weighted evidence groups (combined 100%)Interaction score of Chat Features score
View exact calculation

Each test gets a score from 0 to 10. We multiply every score by how much it counts and add the points together.

Interaction makes up of the Chat Features score. Chat Features makes up 10% of the overall performance score.

Scored tests and weights

Scored testHow much it countsSee scoring
Voice Calls27.00%View
Chat Modes22.00%View
Mode Types17.00%View
Group Chat5.00%View
Double Texting15.00%View
Proactive Messages14.00%View
Total100%

How the score is calculated

  1. Each test gets a score from 0–10

  2. Test score × how much it counts

    Example: 10 × 27.00% = 2.70 points

  3. We do this for every test

  4. We add all the points together

  5. Final Interaction score

    Counts for 30% of Chat Features

Example calculation

We multiply each test score by how much it counts. We then add all the points together.

Scored testTest scoreHow much it countsCalculationPoints added
Voice Calls10.0027.00%10.00 × 27.00%2.70
Chat Modes7.0022.00%7.00 × 22.00%1.54
Mode Types7.5017.00%7.50 × 17.00%1.27
Group Chat5.005.00%5.00 × 5.00%0.25
Double Texting7.0015.00%7.00 × 15.00%1.05
Proactive Messages10.0014.00%10.00 × 14.00%1.40
Final Interaction score100%Add all points8.22/10

Special cases

Not Applicable

If a test does not apply, we remove it and spread its weight across the remaining tests. Mode Types is automatically marked Not Applicable when the app has one chat mode or fewer.

Unknown

If we cannot verify a result, the test receives a score of 0.

Manual adjustment

In rare cases, we may adjust a score when the calculated result is clearly unfair. We always record the reason.

The best AI girlfriend apps do more than wait for you to send a message.

Voice calls let you have a real-time conversation. Chat modes can change how the character behaves. Group chats let several characters take part in the same conversation.

Smaller details also make a big difference.

Double texting feels more natural than receiving one large block of text every time. Proactive messages make the character feel more alive because she can message you without waiting for you to start every conversation.

The difference between apps is huge. One platform may offer calls, several useful chat modes, and characters that message you first. Another may still work like a basic chatbot where every interaction starts and ends with one text reply.

That is why we test whether these features are actually available and whether they work properly.

We use a paid account and test every available interaction feature.

For Voice Calls, we start three calls on three different days. We record whether each call connects and the longest call length the app allows.

For Chat Modes, we count every mode that clearly changes how the conversation works.

We then test two modes with five messages each. We rate whether each mode works well, partly works, or barely changes the chat.

For Group Chat, we create three conversations and try adding two AI characters, three AI characters, and four AI characters.

During normal chat testing, we count how often the AI sends two or more separate messages before we reply.

Finally, we keep three chats open for seven days without sending anything. We record every message the characters send without a new user message.

Some Interaction features may only work with certain characters, devices, or subscription plans.

Voice calls can also be affected by internet speed, microphone permissions, or temporary server problems. Testing on three different days helps reduce the effect of one unusual failure.

Chat modes can be difficult to compare because every platform names them differently. We only count a mode when it clearly changes how the chat behaves.

Double texting does not automatically make the chat better. Sending several useful messages can feel natural, but splitting one basic sentence into five tiny messages can become annoying.

Proactive Messages are tested for seven days. A character may support the feature but message less often than that, so we clearly show the test period with the result.

Our results reflect the paid account, characters, and devices used on the recorded test date.

Evidence groups

Interaction has 6 evidence groups made up of 6 scored tests.

27%

Voice Calls

1 scored test

Voice Calls measures whether you can have a live voice conversation with the AI character.

A proper voice call should feel like a real-time conversation rather than sending separate recorded voice messages.

Voice Calls

Whether you can have a live voice conversation with the AI character.

How we test

We start three voice calls on three different days. For each call, we record whether it connects, whether the audio works, whether the conversation continues normally, and the maximum call length allowed.

What counts
  • A live two-way voice conversation
  • Calls started through the normal chat or call screen
  • The AI responding in real time
  • Calls available to normal paying users
What does not count
  • Recorded voice messages
  • Text replies read aloud
  • Voice previews
  • Pre-recorded audio clips
  • A call button that never connects
Result shown

Yes — all 3 calls connected

Maximum call length: 10 minutes

Scoring

We use the result categories below.

ResultScore
Yes — all three calls worked10/10
Limited — only some calls worked or important restrictions applied5/10
No — voice calls were unavailable or none connected0/10
Unknown0/10

In this example, Voice Calls scores 10/10.

Evidence group 2 of 6

22%

Chat Modes

1 scored test

Chat Modes measures how many different ways you can change how the conversation works.

Examples could include romantic chat, roleplay, storytelling, assistant mode, or other modes that clearly change the AI’s behavior.

Chat Modes

How many different ways you can change how the conversation works.

How we test

We count every selectable mode that creates a noticeable change in the conversation. A different name or icon is not enough. The mode needs to change how the character responds.

What counts
  • Modes that change the reply style
  • Modes that change the type of conversation
  • Story or roleplay modes
  • Modes that add clear rules or behavior
  • Options normal users can select
What does not count
  • Minor tone settings
  • Different names for almost the same mode
  • Buttons that do not noticeably change the replies
  • Character personalities
  • Features that cannot be selected during normal use
Result shown

6 chat modes

Scoring

More working chat modes means a higher score.

Chat modesScore
0 or fewer0/10
13/10
25/10
3–46/10
5–67/10
7–98/10
10+10/10

A result of 6 chat modes scores 7/10.

Evidence group 3 of 6

17%

Mode Types

1 scored test

Mode Types measures whether the available chat modes actually work.

An app can list several modes, but they are not useful when every mode produces almost the same replies.

Mode Types

Whether the available chat modes actually work.

How we test

We select two available modes. We send five messages in each mode and check whether the conversation clearly changes. Each tested mode is rated Good (10 points), Partial (5 points), or Poor (0 points). The Mode Types score is the average of the tested mode ratings.

What counts
  • The mode clearly changes how the AI responds
  • The change stays noticeable across several messages
  • The mode follows its stated purpose
  • The chat remains usable while the mode is active
What does not count
  • A different label with no clear change
  • One unusual reply followed by normal behavior
  • A mode that repeatedly breaks
  • A mode that ignores its own description
  • Character personality settings
Result shown

Average: 7.5/10

Romantic Mode: Good — 10 points. Story Mode: Partial — 5 points.

Scoring

We average the tested mode ratings.

Mode ratingScore
Good10/10
Partial5/10
Poor0/10

In this example, Mode Types scores 7.5/10.

Evidence group 4 of 6

5%

Group Chat

1 scored test

Group Chat measures whether several AI characters can join the same conversation.

This can be useful for group roleplays, stories, or conversations where several characters interact with one another.

Group Chat

Whether several AI characters can join the same conversation.

How we test

We create three group chats. We try adding two AI characters, three AI characters, and four AI characters. We record whether the chats work and the maximum number of characters supported.

What counts
  • Several AI characters inside one conversation
  • Each character clearly identified
  • Characters responding inside the same chat
  • Group chats available to normal users
What does not count
  • Switching between separate one-to-one chats
  • One AI pretending to play several characters
  • A group chat that only includes human users
  • A feature shown in marketing but unavailable in the app
Result shown

Limited — chats with 2 and 3 characters worked, but 4 characters were not supported

Maximum supported: 3 AI characters

Scoring

We use the result categories below.

ResultScore
Yes — all three group-chat tests worked10/10
Limited — only some group sizes worked or important restrictions applied5/10
No — group chat was unavailable0/10
Unknown0/10

In this example, Group Chat scores 5/10.

Evidence group 5 of 6

15%

Double Texting

1 scored test

Double Texting measures how often the AI sends more than one separate message before you reply.

This can make the conversation feel more like a real messaging app instead of receiving one large block of text every time.

Double Texting

How often the AI sends more than one separate message before you reply.

How we test

During normal chat testing, we send a message and wait without replying. We count each time the character sends two or more separate messages before our next message. The result is shown per 100 user messages.

What counts
  • Two or more separate messages sent before the user replies
  • Follow-up messages that add something useful
  • Messages that arrive naturally as part of one reply
What does not count
  • One long message split visually by paragraphs
  • Notifications about credits or app updates
  • Duplicate messages caused by an error
  • Several messages sent only after the user replies again
  • System messages
Result shown

12 double-texting moments per 100 user messages

Scoring

More double-texting moments means a higher score.

Double-texting moments per 100 messagesScore
0 or fewer0/10
1–54/10
6–157/10
16+10/10

A result of 12 double-texting moments scores 7/10.

Evidence group 6 of 6

14%

Proactive Messages

1 scored test

Proactive Messages measures whether the character can message you without waiting for you to write first.

This helps the character feel more present instead of disappearing every time you close the app.

Proactive Messages

Whether the character can message you without waiting for you to write first.

How we test

We keep three active chats open for seven days. We do not send any new messages during the test. We record every message sent by a character without a new user message.

What counts
  • A character starting a new conversation
  • A genuine follow-up to an earlier chat
  • A message sent without a new user prompt
  • Messages delivered through the normal chat
What does not count
  • Marketing notifications
  • Payment reminders
  • System messages
  • Messages scheduled by the user
  • A reply delayed from an earlier user message
  • Push notifications with no message inside the chat
Result shown

Yes — 4 proactive messages arrived during the seven-day test

Messages appeared in 2 of 3 chats

Scoring

We use the result categories below.

ResultScore
Yes — proactive messages worked10/10
Limited — the feature worked with important restrictions5/10
No — no proactive messages arrived or the feature was unavailable0/10
Unknown0/10

In this example, Proactive Messages scores 10/10.

Back to Chat Features