Subscore

QualityTesting methodology

Quality measures how good the finished images actually look.

Most image generators can create one impressive result when everything goes right. We generate a full batch because we want to know whether the app can produce good images consistently—not only one lucky image for its homepage.

We check realism, major visual problems, composition, and the highest resolution you can download.

  • 4 evidence groups
  • 4 scored tests

Quality is organized into 4 evidence groups. Each group contains one scored test: Realism, Visual Errors, Composition, and Resolution. All four tests count equally toward the final Quality score.

Every test gets a score from 0 to 10.

All four tests count equally. We multiply each test score by 25% and add the points together. Quality makes up of the Images score.

  • 25%Realism
  • 25%Visual Errors
  • 25%Composition
  • 25%Resolution
Weighted evidence groups (combined 100%)Quality score of Images score
View exact calculation

Each test gets a score from 0 to 10. We multiply every score by 25% and add the points together.

Quality makes up of the Images score. Images makes up 15% of the overall performance score.

Scored tests and weights

Scored testHow much it countsSee scoring
Realism25.00%View
Visual Errors25.00%View
Composition25.00%View
Resolution25.00%View
Total100%

How the score is calculated

  1. Each test gets a score from 0–10

  2. Test score × how much it counts

    Example: 8.40 × 25.00% = 2.10 points

  3. We do this for every test

  4. We add all the points together

  5. Final Quality score

    Counts for 34% of Images

Example calculation

We multiply each test score by how much it counts. We then add all the points together.

Scored testTest scoreHow much it countsCalculationPoints added
Realism8.4025.00%8.40 × 25.00%2.10
Visual Errors8.0025.00%8.00 × 25.00%2.00
Composition8.2025.00%8.20 × 25.00%2.05
Resolution8.0025.00%8.00 × 25.00%2.00
Final Quality score100%Add all points8.15/10

Special cases

Not Applicable

If a test does not apply, we remove it and spread its weight across the remaining tests.

Unknown

If we cannot verify a result, the test receives a score of 0.

Manual adjustment

In rare cases, we may adjust a score when the calculated result is clearly unfair. We always record the reason.

Images are one of the biggest reasons people sign up for AI girlfriend apps.

The problem is that almost every modern image generator can create something that looks good at first glance. That does not mean it performs well every time.

You may get one great image followed by several with broken hands, damaged faces, strange bodies, or terrible framing.

This becomes even more frustrating when every generation costs credits. A broken image is not only disappointing—it can also mean you need to pay again and hope the next attempt works.

That is why we test a full batch. A strong generator should create good-looking, usable images consistently instead of relying on one lucky result.

We use a paid account and generate 10 test images.

We upload every image to the same worksheet and review them one by one.

For each image, we rate visual realism and overall quality, composition and framing, and major defects or visual problems.

We also download the highest-quality image available and record its exact width and height.

Using a batch of 10 images helps reduce luck. One unusually good or bad result cannot control the whole score.

Image generation always includes some randomness.

The same prompt can create a strong image once and a much weaker result the next time. Testing 10 images reduces this problem, but it cannot remove randomness completely.

Results can also depend on the type of image. A generator may perform well on close-up portraits but struggle with full-body poses, several people, or complicated backgrounds.

Quality only looks at how the finished images appear. Whether the generator followed the prompt and kept the same character is tested separately under Accuracy.

Image models also change quickly. Our results show how the generator performed on the recorded test date.

Evidence groups

Quality has 4 evidence groups made up of 4 scored tests.

25%

Realism

1 scored test

Realism measures how believable and polished the generated images look.

We look at the full image rather than judging only the character’s face.

Realism

How believable and polished the generated images look.

How we test

We rate every image from 1 to 5. When giving the rating, we look at face, body, hands, lighting, and background. A higher rating means the image looks more realistic and has fewer obvious problems.

What counts
  • A believable face
  • Natural-looking body proportions
  • Hands that look complete
  • Lighting that fits the scene
  • A background that looks clear and believable
  • An overall result that feels finished
What does not count
  • Whether the image followed every part of the prompt
  • Whether the same character is preserved across several images
  • Small style choices that are clearly intentional
  • Personal preference for one art style over another
Result shown

Realism result: 84%

Average rating: 4.2 out of 5

Scoring

A higher average rating means a higher score.

Average ratingScore
1 out of 52/10
2 out of 54/10
3 out of 56/10
4 out of 58/10
5 out of 510/10

An average rating of 4.2 out of 5 scores 8.4/10.

The exact average is used between these points.

Evidence group 2 of 4

25%

Visual Errors

1 scored test

Visual Errors measures how many generated images contain at least one major problem.

A generation counts as having an error when the problem is serious enough to make the result look clearly broken or difficult to use.

Visual Errors

How many generated images contain at least one major problem.

How we test

We inspect all 10 images and complete a defect checklist for each one. An image is marked as having a major error when it includes at least one serious problem.

What counts
  • Extra or missing limbs
  • Broken hands
  • A damaged face
  • Objects merged together
  • Broken or disappearing clothing
  • A badly distorted background
  • A body that looks clearly deformed
What does not count
  • A small detail that most people would not notice
  • A style choice that looks intentional
  • A missing prompt detail
  • A weak pose that does not make the image look broken
  • A generation failure that produced no image, which is tested under Experience
Result shown

Visual error rate: 20%

2 of 10 images had a major error

Scoring

Fewer images with major errors means a higher score.

Images with major errorsScore
0%10/10
20%8/10
40%6/10
60%4/10
80%2/10
100%0/10

A visual error rate of 20% scores 8/10.

The exact percentage is used between these points.

Evidence group 3 of 4

25%

Composition

1 scored test

Composition measures how well the subject and background are arranged inside the image.

Even a realistic image can be difficult to use when the character is cut off, placed awkwardly, or surrounded by a messy background.

Composition

How well the subject and background are arranged inside the image.

How we test

We rate every image from 1 to 5. When giving the rating, we look at whether the requested subject is fully visible, accidental cropping, subject placement, background clarity, and overall balance.

What counts
  • The character is visible as requested
  • Important body parts are not accidentally cropped
  • The subject is placed naturally
  • The background is clear
  • The full image feels balanced
What does not count
  • Whether the prompt details are correct
  • Small creative framing choices
  • Personal preference for close-up or full-body images
  • Visual defects already counted under Visual Errors
Result shown

Composition result: 82%

Average rating: 4.1 out of 5

Scoring

A higher average rating means a higher score.

Average ratingScore
1 out of 52/10
2 out of 54/10
3 out of 56/10
4 out of 58/10
5 out of 510/10

An average rating of 4.1 out of 5 scores 8.2/10.

The exact average is used between these points.

Evidence group 4 of 4

25%

Resolution

1 scored test

Resolution measures the maximum image size you can download.

A larger image gives you more detail and is easier to crop, edit, or use on a larger screen.

Resolution

The maximum image size you can download.

How we test

We generate an image using the highest quality setting available. We download the finished file and record its exact width and height in pixels. We use the downloaded file rather than trusting the resolution shown on the pricing or marketing page.

What counts
  • The highest-quality image normal users can download
  • The actual width and height of the finished file
  • Upscaling that is included and available through normal use
What does not count
  • Resolution promised only in marketing
  • A preview image that cannot be downloaded
  • A larger display size that does not change the real file
  • Third-party upscaling outside the app
Result shown

Maximum resolution: 1920 × 1080

Resolution level: 1080p

Scoring

We use the ranges below.

Maximum resolutionScore
480p4/10
720p6/10
1080p8/10
4K10/10

A maximum resolution of 1080p scores 8/10.

Back to Images