How our rating system works
Our rating system has four levels. Specific evidence is used to calculate subscores, which then determine each category rating. Finally, the weighted category ratings are combined to create the overall performance score.
Scores throughout the site include interactive explanations.
Overall performance score
The overall performance score is the final rating shown across the site. It gives you a quick overview of how well an app performs across all eight categories.
It is not a simple average. Each category has a different weight based on how important it is to the overall AI companion experience.
The box on the right shows an example of what an overall performance score looks like.
Our eight testing categories
Every app receives a score across eight main categories. Each category is weighted based on how important it is to the overall AI companion experience.
Characters
View Characters methodologyMeasures the platform's ready-made character library.
Customization
View Customization methodologyMeasures how much control users have when creating their own character.
Measures the quality of the actual conversation.
Chat Features
View Chat Features methodologyMeasures what users can do inside the chat.
Images
View Images methodologyMeasures image quality and the image-generation experience.
Measures video capabilities, output quality and generation experience.
Privacy
View Privacy methodologyMeasures chat privacy, user control, account security, billing privacy and support.
Pricing
View Pricing methodologyMeasures what users pay and the value they receive.
You are probably wondering how we came up with them.
Herman, the CEO of AI Girlfriend Expert, was a beta tester for some of the first AI girlfriend apps back in October 2024, including DreamGF AI, Candy AI, Character AI, and Replika.
At the time, chat was the main feature available because AI image generation was still very new and not very good. This forced many AI girlfriend apps to build more features around chatting, such as interactive roleplay scenarios, the ability to upload memory snippets, and even AI group chats.
That is how our Chat and Chat Features categories were born.
Privacy became a category because it is important for any website, especially one with adult content.
Pricing is also included because you want to know whether you are getting good value for your money.
The other four categories were developed through extensive research and a survey of 100 AI girlfriend users across different platforms. We asked them which features mattered most. We found participants by joining more than 25 Discord servers from different AI girlfriend platforms.
How subscores work
Each category is divided into three broad subscores. These organize related measurements into results that are easier to understand and compare.
Example: Chat
View Chat methodologyThe Chat category is evaluated across 3 subscores.
Understanding
How well the AI understands the user: memory, relevance, multi-message context, instruction following and roleplay accuracy.
View Understanding methodologyRealism
How natural and human the conversation feels: naturalness, personality consistency, roleplay quality, initiative, emotion and style.
View Realism methodologyReliability
Technical dependability of the chat: repetition, unnecessary refusals, reply speed, errors, consistency and recovery from misunderstandings.
View Reliability methodologyEvidence groups
Evidence groups are the individual measurements collected during testing. They are the foundation of every score. Some are simple yes/no checks; others use percentages, counts, timings, or structured quality assessments.
Memory
We give the app a fixed set of facts during conversations and later test how many it recalls correctly.
Displayed result: 82% of tested facts remembered
View Memory test methodologyReply speed
We record response times across a fixed sample and calculate the median.
Displayed result: 3.8-second median response time
View Reply speed test methodologyCharacter consistency
We generate repeated images of the same character and measure how consistently identity and appearance are preserved.
Displayed result: 84% character consistency
View Character consistency test methodologyExplore our full testing framework
10%CharactersView Characters methodology
34%VarietyView Variety methodology
| Amount | Styles |
| Genders | Ethnicities |
| Personalities | Scenarios |
33%DiscoveryView Discovery methodology
| Filters | Categories |
| Search | Browsing |
33%QualityView Quality methodology
| Duplicates | Originality |
| Profile Quality | Visual Quality |
15%CustomizationView Customization methodology
34%AppearanceView Appearance methodology
33%PersonalityView Personality methodology
| Traits | Interests |
| Relationship | Role |
| Voice | Kink Options |
33%ControlView Control methodology
| Custom Prompts | Editing |
| Preview |
20%ChatView Chat methodology
34%UnderstandingView Understanding methodology
| Memory | Relevance |
| Context | Instructions |
| Roleplay Accuracy |
33%RealismView Realism methodology
| Naturalness | Personality |
| Roleplay | Initiative |
| Emotion | Style |
33%ReliabilityView Reliability methodology
| Repetition | Refusals |
| Reply Speed | Errors |
| Consistency | Recovery |
10%Chat FeaturesView Chat Features methodology
30%MediaView Media methodology
30%InteractionView Interaction methodology
10%Platform ExtrasView Platform Extras methodology
| Live Cam | Other Extras |
15%ImagesView Images methodology
34%QualityView Quality methodology
| Realism | Visual Errors |
| Composition | Resolution |
33%ExperienceView Experience methodology
10%VideoView Video methodology
34%CapabilitiesView Capabilities methodology
33%QualityView Quality methodology
33%ExperienceView Experience methodology
| Speed | Failures |
| Ease of Use | Regeneration |
10%PrivacyView Privacy methodology
31%Data UseView Data Use methodology
| Training | Human Review |
| Data Sharing | Advertising |
| Retention | Policy Clarity |
28%User ControlView User Control methodology
28%SecurityView Security methodology
13%SupportView Support methodology
10%PricingView Pricing methodology
30%Plan ValueView Plan Value methodology
35%Usage CostsView Usage Costs methodology
| Image Cost | Video Cost |
| Voice Cost | Call Cost |
| Top-Up Value | Monthly Spend |
20%Free AccessView Free Access methodology
15%BillingView Billing methodology
How we test AI girlfriend apps in practice
We purchase access
We buy the paid plans ourselves and test the same versions available to regular customers. We do not rely on sponsored access or limited demos. We go through the same customer journey as everyone else.
We create a test plan
Each app is tested using the same category framework, structured tasks, and fixed evidence requirements.
We use the app like a real customer
We create characters, hold conversations, generate media, including NSFW content, test account controls, and examine the real cost of regular use.
We collect evidence
We record counts, percentages, timings, feature availability, failures, limits, and qualitative test results.
We calculate the scores
Scored tests contribute to subscores. Subscores create category scores, and the weighted category scores produce the overall performance score.
We write and fact-check the review
Our editorial conclusions explain the data rather than replace it.
We include personal experience and a video review
One of our experts creates a detailed review, which may also include a YouTube video. This does not contribute to the performance rating, but it helps explain our scores.
We retest and update
We update scores when pricing, features, models, policies, or output quality change in a meaningful way. We check every review at least once every three months, but many are updated more often because the industry changes quickly.
Where you will see our scores
- Reviews
Full overall, category, subscore, and evidence-level results for one app.
- Roundups
Condensed results used to compare and rank several apps.
- Comparisons
Side-by-side category and evidence comparisons.
- App directory
Summary scores, pricing information, and key strengths for browsing and filtering.
Beyond the score
We aim to keep our reviews as objective as possible, so our recommendations are truly based on data. That said, personal experience, expert analysis, and how a platform actually feels to use still matter. We are humans, after all.
- Market data methodology
Google Trends, search interest, growth, internal popularity rankings, and competitive comparisons.
- Editorial guidelines
Independence, fact-checking, corrections, and affiliate disclosures.
Scored vs informational
| Content type | Affects score |
|---|---|
| Structured category tests | Yes |
| Pricing evidence | Yes |
| Personal written review | No |
| YouTube video review | No |
| Photos and videos gallery — Supporting evidence only | No |
| Google Trends and market data | No |
| Expert analysis and key takeaways | No |
How we keep scores consistent
Every app is tested using the same categories and the same tasks.
When privacy information is unclear or unavailable, we do not assume the app is safe.
Image and video costs are shown in those sections, but they only affect the Pricing score. This prevents us from counting the same thing twice.
Simply having a feature does not guarantee a high score. The feature also needs to work well.
Updates and methodology versions
Current methodology: V3.1 (August 2026). We recalculate or retest when products or methodology versions change.
| Version | Date | Main change |
|---|---|---|
| V3.1 | August 2026 | Updated image and pricing tests |
| V3.0 | July 2026 | Introduced eight-category framework |
| V2 | Early 2025 | Expanded categories using a survey of AI girlfriend users |
| V1 | October 2024 | Initial ratings based on personal experience |
Frequently asked questions
Real questions. Honest answers.
Can companies pay for a higher score?
No. Companies cannot buy a better rating, change their test results, or pay to appear higher in our rankings. Scores come from the same testing method for every app. We may earn a commission when someone uses one of our links, but this does not affect the score.
Do you test free or paid accounts?
We mainly test paid accounts because they show what a regular customer receives after subscribing. We also test free access separately when an app offers a free plan or trial.
How long do you test each app?
We normally use an app for at least 30 days before publishing its score. This gives us enough time to test features such as memory, proactive messages, customer support, and account controls that cannot be judged in one day.
How often do you update scores?
We update a score when an important part of the app changes. This may include its pricing, features, AI model, privacy policy, limits, or image and video quality. We also retest products over time to check whether their performance has improved or become worse.
Why do some apps have no video score?
Some apps do not offer video generation. In that case, we mark Video as not available instead of giving the app a made-up score. The remaining available categories are then used to calculate the overall result.
How is the overall score calculated?
The overall score combines eight categories:
- Characters
- Customization
- Chat
- Chat Features
- Images
- Video
- Privacy
- Pricing
Each category has a different weight, so the overall score is not a simple average. More important categories have a larger effect on the final result.
Are any scores based on opinion?
Some tests require human judgment, especially when we rate realism, writing quality, image errors, or video motion. To keep this fair, we use fixed sample sizes, the same testing steps, and written scoring rules for every app.
What happens when information is unclear or unknown?
We never treat missing information as proof that an app is safe or trustworthy. When we cannot confirm something, we mark it as Unknown and explain what information was missing.
Can an app’s score go down?
Yes. A score can decrease when retesting shows worse performance, prices increase, features are removed, policies become less clear, or our methodology is improved.
How do you avoid counting the same feature twice?
We may test the same feature in more than one place because it can answer different questions. For example, we record image and video costs in their feature sections, but those costs only affect the final score under Pricing. This prevents one result from unfairly adding points or penalties more than once.

