Skip to content

How we test

Every score has a test behind it.

One inbox of 412 real messages, the same rubric, the same reviewer, published side by side with the raw notes. Where we haven't tested yet, we say “Not yet tested” instead of guessing.

The shared test inbox

Every tool is connected to the same test inbox: 412 real messages spanning newsletters, cold outreach, receipts, scheduling threads, and conversations awaiting a reply. Each tool sees identical mail, so differences in behavior are differences between tools — not differences in input.

The same reviewer runs every test, works through the same task list (triage the backlog, answer the threads that need answers, unsubscribe from the junk), and keeps raw notes that are published with the score.

The five scoring dimensions

Each tested tool is scored 0–10 on five dimensions, weighted equally at launch. Use-case pages that reweight dimensions will publish their weights alongside the ranking.

Dimension What we evaluate
Triage accuracy Did the tool label, prioritize, and file the inbox the way a careful assistant would? We count misses: important mail buried, newsletters surfaced, cold outreach let through.
Draft quality How much editing did generated replies need? We compare drafts against replies we actually sent, looking at tone match, factual grounding in the thread, and how often a draft was usable as-is.
Setup effort Time from signup to useful behavior, the number of decisions forced on the user, and how recoverable mistakes are during onboarding.
Privacy posture What mailbox permissions are requested, what leaves your mailbox, published data-handling and retention terms, and whether self-hosting or open source reduces exposure.
Price for value Price against what the tool actually did in our test — not against its feature list. Free tiers and trials count when they are genuinely usable.

Rules we hold ourselves to

  • No score without a test. Untested tools are listed with verifiable facts only and are excluded from ranked positions.
  • Failures are published. Every profile names at least one real limitation before it links to the vendor.
  • Placeholders are labeled. Any price, claim, or score that has not been verified carries a visible “Placeholder” marker. We never present invented numbers as fact.
  • Facts carry dates. Pricing and features carry a last-verified date; stale entries lose their score until re-verified.
  • Evidence is linked. Scores link to the test notes and screenshots that produced them.

Conflicts of interest

This site is owned by Inbox Zero Inc., maker of Inbox Zero — one of the tools in the directory. Inbox Zero is tested with the same inbox, the same rubric, and the same reviewer as everything else, and its limitations are published alongside its score. Sponsorship and submissions never affect editorial scores; the full rules live in the editorial policy →

Current status

The directory is pre-launch. Dimension scores and several facts shown across the site are representative placeholders carried from the design prototype, labeled as such, and will be replaced by the first published test cycle before anything is presented as verified.

Built something we should test?

Submissions go to a moderation queue, not straight to a public listing. Nothing gets a score until we've run the inbox against it ourselves.