guide

How We Test AI Story Tools (And What We Do Not Test)

We make one of the tools we compare. This page explains how the comparisons are built, what we measure, what we refuse to guess at, and how to correct us.

ReadKidz Team 7 min read

Key points

  1. We make one of the tools we compare

    ReadKidz publishes these comparisons and is one of the products in them, which is a conflict of interest that is better stated than hidden.
  2. Six dimensions, five of them checkable from outside

    Language coverage, format breadth, free tier, entry price and printed-book delivery can all be verified from a vendor own public pages; character consistency cannot, so it is scored only where we have run the tool.
  3. What we do not test, and why we say so

    We have not generated a book in every tool we list, so illustration and character consistency are left blank for tools we have not run — a blank is not a low score.
  4. What we will not do

    We do not invent numbers, invent authors, or publish self-assigned star ratings as structured data.

The short version

We publish comparisons of AI story tools, and we make one of the tools in them. That is a conflict of interest, and it is better stated than hidden.

What makes such a comparison worth reading is not neutrality we cannot claim. It is that the method is public and the numbers are checkable. This page is the method.

What we measure

Six dimensions, applied identically to every tool including our own:

DimensionHow it is established
Language coverageCounted from the language versions a site actually publishes, not from marketing claims. Where only a vendor claim exists, it is labelled as one
Output format breadthPicture book, comic, animation or video, audio narration, long-form, open image generation
Free tierWhether one exists, whether the output is usable, whether a card is required
Entry priceThe lowest published price that produces a finished book
Printed-book deliveryNone, print-ready file only, or printed and shipped
Character consistencyWhether one character keeps the same hair, clothes and face across eight or more pages

The first five can be verified by anyone with a browser. The sixth cannot.

Where the numbers come from

Every figure is read off the vendor own public pages, and each article prints the date it was captured. The current set was captured on 31 August 2026.

Three rules govern what gets printed:

  1. A published number beats a claimed one. Language coverage is counted from the language versions a site serves, because publishing a Korean version is a commitment in a way that a homepage sentence is not.
  2. Where only a claim exists, it is labelled a claim. One tool in our current comparison states eleven languages without publishing them; the table says "vendor stated".
  3. Undisclosed is written as undisclosed. Several tools do not publish pricing. We write "not disclosed" rather than estimating, even where estimating would make the table look more complete.

What we do not test

This is the section other comparisons leave out, and it is the one that decides whether the rest is worth anything.

We have not generated a book in every tool we list. Illustration quality and character consistency need hands-on use, so those are scored only for tools we have actually run and left blank for the others. A blank is not a low score. Publishing a score we did not earn would quietly devalue every number beside it.

We have not benchmarked prose quality per language, for anyone, including ourselves. We can tell you which tools publish in Thai. We cannot tell you whose Thai reads better, and we do not pretend to.

We do not measure how tools behave over time — whether a free tier shrinks, whether print quality holds. Those need months of use, not an afternoon of checking.

How we handle our own product

ReadKidz is scored on the same six dimensions, and the results are printed whether or not they flatter us. In the current set that means saying, in our own articles, that:

Each article also opens with a table of situations, and several rows recommend a competitor.

What we will not do

Three practices are common in this category. We do not use them, and it is worth being explicit about why.

We do not invent authors. Some comparison sites publish under a reviewer name with a portrait, where no such person exists. That is a fabricated identity. We publish under the organisation name instead — less warm, and true.

We do not publish self-assigned ratings as structured data. Emitting your own star rating for search engines to display breaches their policy, and the penalty lands on the whole site rather than the one page.

We do not refresh dates to look current. One tool in our current comparison has stamped 96% of its pages with the same recent date. Republishing an unchanged page with a new date is not maintenance; re-checking the facts is, and that is why our figures carry a capture date rather than a freshness badge.

Keeping it current

Pricing and feature claims in this category change on other people schedules. Our standard is that no competitor figure older than 120 days should stand without being re-checked, and that a re-check means opening the vendor page again — not editing the date.

Where a figure has changed since we published, we change it and say that it changed.

Tell us if we got something wrong

If a vendor page shows something different from what we published, we would rather hear it than not. Email sales@readkidz.com with the page and the correction.

This applies most of all to the tools we compare ourselves against. A comparison written by a competitor earns its place by being correctable in public.

Key data

Dimensions scored in a ReadKidz comparison
6
Source: ReadKidz editorial rubric · 2026
Dimensions verifiable from vendor public pages
5 of 6
Source: ReadKidz editorial rubric · 2026
Capture date of the current competitor figures
31 August 2026
Source: ReadKidz competitor fact table · 2026
Maximum age of a competitor figure before we re-check it
120 days
Source: ReadKidz editorial standard · 2026

Where ReadKidz fits

Questions parents ask

Who writes these comparisons?

The ReadKidz team, in-house. We publish under the organisation name rather than inventing an individual reviewer persona, and we state that we make one of the products being compared.

How can a comparison written by a competitor be fair?

By being checkable. Every figure names its source and its capture date, every score uses the same published rubric, and each article states the situations where a different tool is the better choice. If a claim is wrong, it can be checked against the vendor own page in a minute.

Do you test every tool you write about?

No, and we say which ones. Pricing, languages, formats and print options are read from public pages. Illustration and character consistency need hands-on generation, so they are scored only for tools we have run and left blank elsewhere.

How current are the figures?

Each article prints the capture date. Pricing in this category changes often, so we re-check figures rather than letting them age quietly, and we would rather you confirm on the vendor own site before buying.

You got something wrong. How do I tell you?

Email sales@readkidz.com with the page and the correction. If a vendor page shows something different from what we published, we change it and note that it changed.