# How the Index is made

What is collected, how it is filtered and scored, how much the numbers can be trusted, and where the method falls short today.

Method v0.2, published September 17, 2026 · Data Jun 19 to Sep 16, 2026

## What counts as a dictation tool

Software that types what you say into the application or text field you are already using. That includes AI dictation apps, dictation built into an operating system, browser extensions and web notepads, open-source dictation, voice-control systems that include dictation, and professional and medical dictation.

Meeting recorders, file transcription, raw speech-to-text APIs and text-to-speech are out of scope. Listing is automatic for anything that meets the definition. 69 tools are in the [registry](https://dictationindex.com/registry) today.

## From posts to a score

### 1. Collect

Public posts on X that mention a tool by name, inside a rolling 90-day window. Collection stops at 1,000 posts per tool, newest first, so the most discussed tools cover fewer days than the rest. See date coverage below.

### 2. Classify

Each post is labelled positive, negative, neutral or mixed toward the tool. It is also labelled with the topics it discusses (accuracy, speed, price, reliability, privacy, platform support, support and usability), each with its own polarity, and with the usage context the author describes.

### 3. Filter

Posts that are not about the tool, duplicates and posts that could not be labelled are removed, along with official vendor accounts and automated reply bots. No single author can contribute more than three posts to one tool: their three most recent are kept. What is left is n, the number of qualifying posts.

### 4. Score

Each post counts 10 if positive, 5 if neutral or mixed, and 0 if negative. The average is pulled toward the midpoint of the scale by a small fixed prior, so a handful of glowing posts cannot outrank a large, mostly positive sample.

`score = (10 × pos + 5 × neu + m × μ) / (n + m)`

The prior mean μ is fixed at 5 and the prior weight m at 5 posts. Scores are rounded to one decimal. You can try the formula with your own numbers on the [homepage calculator](https://dictationindex.com/#method).

## Margins, ties and the minimum sample

**Margin.** Every ranked score is printed with a ± figure, which is a 90% margin of error. Small samples have wide margins: VoiceOS, with 80 posts, carries ± 0.5, while Typeless, with 769, carries ± 0.2.

**Ties.** Where two scores are closer than their margins allow, the tools are statistically tied and every ranking row says which ranks a tool is tied with. In the current edition Aqua Voice, Spokenly, Handy, VoiceOS, OpenWhispr, Willow Voice, VoiceInk and MacWhisper form one leading group. Inside a tied group the order is not settled, and a rank number should not be quoted without its tie.

**Minimum sample.** A tool needs 50 qualifying posts to be ranked. 12 tools clear that today. Tools below it are listed with their facts and their post count, without a rank or a score. A tool whose collected posts are mostly about something else that shares its name is not scored at all.

**Category lists.** The [best-of lists](https://dictationindex.com/best) reuse the same score, filtered by registry facts such as platform or licence. A list exists only when at least three of its tools are ranked.

## Date coverage

The window is 90 days, but the 1,000-post cap means the 4 most discussed tools are scored on a shorter, more recent stretch. A few days of posts can reflect one news cycle instead of settled opinion, so treat the shortest spans with the most care. Every ranking row and tool page prints the dates its posts cover.

**Ranked tools whose posts cover less than the full window**

| Tool | Posts | Posts cover | Days |
| --- | --- | --- | --- |
| [Aqua Voice](https://dictationindex.com/tools/aqua-voice) | 606 | Aug 5 to Sep 16, 2026 | 43 |
| [Superwhisper](https://dictationindex.com/tools/superwhisper) | 554 | Aug 21 to Sep 16, 2026 | 27 |
| [Typeless](https://dictationindex.com/tools/typeless) | 769 | Aug 22 to Sep 16, 2026 | 26 |
| [Wispr Flow](https://dictationindex.com/tools/wispr-flow) | 668 | Sep 11 to Sep 16, 2026 | 6 |

The fix is to sample evenly across the whole window instead of taking the newest 1,000 posts. It is planned for the next method version.

## Known limits

These are the reasons to hold the scores loosely. None of them is hidden, and each is a commitment to fix or to keep disclosing.

### The method embeds choices

The source, the filters, the 50-post minimum, the prior and the formula are all judgment calls. Different reasonable choices would move some tools, most of all inside a tied group.

### The counts cannot be re-run yet

The post IDs and the search terms behind each score are not published. Until they are, nobody outside can check a count, so every score is our count. Publishing a per-tool file of post IDs with their labels is the next thing we owe readers.

### The labelling step is not documented

How each post gets its sentiment and topic labels, the instructions used, and the accuracy of those labels against a human-labelled sample have not been published. Sarcasm, jokes and posts that compare two tools are the likeliest places for a wrong label.

### One source, with a lean

Every post comes from X. People who post there about dictation are disproportionately software builders and early adopters, and they mostly write in English. Medical, legal and accessibility tools are barely discussed there, which is why 57 of 69 registry tools have no score. Reddit is planned as the second source. App store reviews and review sites are not included.

### Affiliated posters are not filtered

Official vendor accounts are removed. Posts from employees, founders, investors, affiliates and paid creators using personal accounts are not yet identified, for any tool. The three-posts-per-author cap limits the damage one account can do, but not a coordinated push.

### Neutral posts pull scores toward the middle

A post that mentions a tool without an opinion counts 5. Tools that attract a lot of neutral chatter therefore sit closer to 5 than tools discussed mostly by people with a view. The sentiment split on each tool page shows how much of a score is neutral.

### Uneven date coverage

Covered above: the most discussed tools are scored on days or weeks of posts, the rest on up to 90 days.

## Versions and editions

The method has a version number. Any change to collection, labelling, filtering or scoring raises it. Each published ranking is kept as a dated [edition](https://dictationindex.com/editions) with a permanent address and the method version it used, so a citation keeps pointing at the numbers it quoted.

Found a mistake in the method or in a tool's facts? [Tell us](https://dictationindex.com/corrections).

---

Source: https://dictationindex.com/method
The Dictation Index, September 2026 edition. Data through 2026-09-16. Method v0.2: https://dictationindex.com/method
Scores are aggregated sentiment from public posts on X. Quote a rank together with its margin and any statistical tie.
All data as JSON (CC BY 4.0): https://dictationindex.com/tools.json · Everything in one file: https://dictationindex.com/llms-full.txt
