How We Test AI Girlfriend AppsTesting methodology

Recent updates

Herman Carter
Lead Reviewer

In a world where everyone hammers out AI-generated affiliate articles and calls themselves an “expert” who gives honest reviews, we do things differently.

Every AI girlfriend app mentioned on aigirlfriend.expert goes through a rigorous review process that relies heavily on data-driven scores.

This page explains our entire scoring process, what those scores mean, and how we calculate them.

  • 1product reviewed
  • 8rating categories
  • 27subscores
  • 153scored tests

How our rating system works

Our rating system has four levels. Specific evidence is used to calculate subscores, which then determine each category rating. Finally, the weighted category ratings are combined to create the overall performance score.

Scores throughout the site include interactive explanations.

Overall performance score

The overall performance score is the final rating shown across the site. It gives you a quick overview of how well an app performs across all eight categories.

It is not a simple average. Each category has a different weight based on how important it is to the overall AI companion experience.

The box on the right shows an example of what an overall performance score looks like.

Our eight testing categories

Every app receives a score across eight main categories. Each category is weighted based on how important it is to the overall AI companion experience.

You are probably wondering how we came up with them.

Herman, the CEO of AI Girlfriend Expert, was a beta tester for some of the first AI girlfriend apps back in October 2024, including DreamGF AI, Candy AI, Character AI, and Replika.

At the time, chat was the main feature available because AI image generation was still very new and not very good. This forced many AI girlfriend apps to build more features around chatting, such as interactive roleplay scenarios, the ability to upload memory snippets, and even AI group chats.

That is how our Chat and Chat Features categories were born.

Privacy became a category because it is important for any website, especially one with adult content.

Pricing is also included because you want to know whether you are getting good value for your money.

The other four categories were developed through extensive research and a survey of 100 AI girlfriend users across different platforms. We asked them which features mattered most. We found participants by joining more than 25 Discord servers from different AI girlfriend platforms.

How subscores work

Each category is divided into three broad subscores. These organize related measurements into results that are easier to understand and compare.


20%

The Chat category is evaluated across 3 subscores.

Realism

How natural and human the conversation feels: naturalness, personality consistency, roleplay quality, initiative, emotion and style.

View Realism methodology

Reliability

Technical dependability of the chat: repetition, unnecessary refusals, reply speed, errors, consistency and recovery from misunderstandings.

View Reliability methodology

Evidence groups

Evidence groups are the individual measurements collected during testing. They are the foundation of every score. Some are simple yes/no checks; others use percentages, counts, timings, or structured quality assessments.

Memory

We give the app a fixed set of facts during conversations and later test how many it recalls correctly.

Displayed result: 82% of tested facts remembered

View Memory test methodology

Explore our full testing framework

10%CharactersView Characters methodology
15%CustomizationView Customization methodology
20%ChatView Chat methodology
10%Chat FeaturesView Chat Features methodology
15%ImagesView Images methodology
10%VideoView Video methodology
10%PrivacyView Privacy methodology
10%PricingView Pricing methodology

How we test AI girlfriend apps in practice

01

We purchase access

We buy the paid plans ourselves and test the same versions available to regular customers. We do not rely on sponsored access or limited demos. We go through the same customer journey as everyone else.

02

We create a test plan

Each app is tested using the same category framework, structured tasks, and fixed evidence requirements.

03

We use the app like a real customer

We create characters, hold conversations, generate media, including NSFW content, test account controls, and examine the real cost of regular use.

04

We collect evidence

We record counts, percentages, timings, feature availability, failures, limits, and qualitative test results.

05

We calculate the scores

Scored tests contribute to subscores. Subscores create category scores, and the weighted category scores produce the overall performance score.

06

We write and fact-check the review

Our editorial conclusions explain the data rather than replace it.

07

We include personal experience and a video review

One of our experts creates a detailed review, which may also include a YouTube video. This does not contribute to the performance rating, but it helps explain our scores.

08

We retest and update

We update scores when pricing, features, models, policies, or output quality change in a meaningful way. We check every review at least once every three months, but many are updated more often because the industry changes quickly.

Where you will see our scores

  • Reviews

    Full overall, category, subscore, and evidence-level results for one app.

  • Roundups

    Condensed results used to compare and rank several apps.

  • Comparisons

    Side-by-side category and evidence comparisons.

  • App directory

    Summary scores, pricing information, and key strengths for browsing and filtering.

Beyond the score

We aim to keep our reviews as objective as possible, so our recommendations are truly based on data. That said, personal experience, expert analysis, and how a platform actually feels to use still matter. We are humans, after all.

Scored vs informational

Content typeAffects score
Structured category testsYes
Pricing evidenceYes
Personal written reviewNo
YouTube video reviewNo
Photos and videos gallery — Supporting evidence onlyNo
Google Trends and market dataNo
Expert analysis and key takeawaysNo

How we keep scores consistent

Every app is tested using the same categories and the same tasks.

When privacy information is unclear or unavailable, we do not assume the app is safe.

Image and video costs are shown in those sections, but they only affect the Pricing score. This prevents us from counting the same thing twice.

Simply having a feature does not guarantee a high score. The feature also needs to work well.

Updates and methodology versions

Current methodology: V3.1 (August 2026). We recalculate or retest when products or methodology versions change.

VersionDateMain change
V3.1August 2026Updated image and pricing tests
V3.0July 2026Introduced eight-category framework
V2Early 2025Expanded categories using a survey of AI girlfriend users
V1October 2024Initial ratings based on personal experience

Frequently asked questions

Real questions. Honest answers.

Can companies pay for a higher score?

No. Companies cannot buy a better rating, change their test results, or pay to appear higher in our rankings. Scores come from the same testing method for every app. We may earn a commission when someone uses one of our links, but this does not affect the score.

Herman Carter
Do you test free or paid accounts?

We mainly test paid accounts because they show what a regular customer receives after subscribing. We also test free access separately when an app offers a free plan or trial.

Herman Carter
How long do you test each app?

We normally use an app for at least 30 days before publishing its score. This gives us enough time to test features such as memory, proactive messages, customer support, and account controls that cannot be judged in one day.

Herman Carter
How often do you update scores?

We update a score when an important part of the app changes. This may include its pricing, features, AI model, privacy policy, limits, or image and video quality. We also retest products over time to check whether their performance has improved or become worse.

Herman Carter
Why do some apps have no video score?

Some apps do not offer video generation. In that case, we mark Video as not available instead of giving the app a made-up score. The remaining available categories are then used to calculate the overall result.

Herman Carter
How is the overall score calculated?

The overall score combines eight categories:

  • Characters
  • Customization
  • Chat
  • Chat Features
  • Images
  • Video
  • Privacy
  • Pricing

Each category has a different weight, so the overall score is not a simple average. More important categories have a larger effect on the final result.

Herman Carter
Are any scores based on opinion?

Some tests require human judgment, especially when we rate realism, writing quality, image errors, or video motion. To keep this fair, we use fixed sample sizes, the same testing steps, and written scoring rules for every app.

Herman Carter
What happens when information is unclear or unknown?

We never treat missing information as proof that an app is safe or trustworthy. When we cannot confirm something, we mark it as Unknown and explain what information was missing.

Herman Carter
Can an app’s score go down?

Yes. A score can decrease when retesting shows worse performance, prices increase, features are removed, policies become less clear, or our methodology is improved.

Herman Carter
How do you avoid counting the same feature twice?

We may test the same feature in more than one place because it can answer different questions. For example, we record image and video costs in their feature sections, but those costs only affect the final score under Pricing. This prevents one result from unfairly adding points or penalties more than once.

Herman Carter