# Xero billing guide: how to compare AI-generated and editor-tested drafts

Brady Weaver · October 5, 2026

> Compare AI-generated and editor-tested Xero billing drafts using observable actions, single-page clarity, and practical checks that help users act confidently.

Before comparing anything, convert each section into an observable action. Ask what the reader does after reading it: locate a billing setting, complete a step, interpret an error response, confirm that the account state changed. A section that maps to no action cannot be scored. A documentation-engineering discussion on LinkedIn frames a related test — whether a non-technical user can understand an AI product's limitations without clicking to a second page — and the same single-page rule applies here. If the reader must open another page to name the action, the section has not finished the sequence.

Hold the order fixed: question, packet, draft, action, verification. Nothing advances because a draft was produced quickly. *The Washington Post* reviewed five AI bots on one tough reading test and found the strongest performer was not the best-known model — the shared test, not reputation, produced the signal. The pipeline ends when an editor re-verifies every Xero billing claim against the manifest before publish; each mapped action returns a result along the way, and those results are what the comparison later reads.

![Before comparing anything, convert each section into an — Xero billing guide](https://static.mm-ais.com/article-images-ai/xero-billing-guide-how-to-compare-ai-gen-ai-73f57841.jpg)

## Build an evidence convergence check

An evidence convergence check runs before any Xero billing comparison reaches a reader, and its first result is negative: the supplied source set contains no Xero billing comparison study, no participant result, and no comprehension percentage. Adjacent material does not fill the gap. The Washington Post's five-bot reading test scores general reading, and the document-comprehension directory on There's An AI For That catalogs tools rather than testing Xero billing tasks. Until a new test produces figures, state no observed winner, and tag every unmeasured value — first-pass task success, correction events, time on task — as a method target, never as a finding.

Ground each factual claim in three named records: the applicable Xero Central article, the applicable Xero Developer Documentation page, and the dated test log or repository revision. Naming is literal and checkable — article title or URL, documentation page path, and the log file with its date or the commit identifier. A claim backed by one or two records stays a candidate; it does not enter the comparison.

When the three records disagree, preserve the disagreement rather than resolving it in prose. Put the conflicting statements side by side with their sources, mark the claim unresolved, and route it to an editor for a ruling and a re-check date. Averaging two sources, or quoting only the one that matches the draft, hides the conflict the check exists to surface.

| Record | What it can verify | Action when it is the only source |
| --- | --- | --- |
| Xero Central article | User-facing billing behavior and terminology | Hold the claim as a candidate until the other records arrive |
| Xero Developer Documentation page | API parameters and payload behavior behind the billing feature | Hold; a documented parameter is not a measured reader outcome |
| Dated test log or repository revision | What a specific draft caused a specific reader to do, tied to a date or commit | Hold; one run is one record, not convergence |

Tag each sentence that carries a number: Verified when three records agree, Unresolved when they conflict, Target when no measurement exists yet. A sentence with no tag and no source gets deleted, not softened. The tags keep the section auditable in both directions: a reader can trace a figure back to a record, and an editor can see which claims still need work.

Nothing publishes until an editor re-verifies every Xero billing claim against the three records again, including claims cleared in an earlier pass. That second pass comes after the final edit, because late edits introduce new claims and quietly invalidate earlier checks.

![Build an evidence convergence check — Xero billing guide](https://static.mm-ais.com/article-images-ai/xero-billing-guide-how-to-compare-ai-gen-ai-02aa7dc1.jpg)

## Compare drafts with one explicit winner

The decision matrix below is the only artifact this section adds, and it should be filled in before anyone argues about which Xero billing draft reads better. Every row takes the same measurement from both drafts — the AI-generated one and the editor-tested one — and every row carries a winner rule that was fixed in advance, so the outcome is settled by entries instead of preference or volume. Enter observed values only; leave a cell blank rather than estimate it. Drafting speed has no row here.

| Criterion | AI-generated draft | Editor-tested draft | Winner rule |
| --- | --- | --- | --- |
| First-pass task success | Record observed count and denominator | Record observed count and denominator | Higher rate wins |
| Correction events | Count wrong, missing, or reversed actions | Count wrong, missing, or reversed actions | Lower count wins if the rates tie |
| Source traceability | Mark each claim as linked or unlinked | Mark each claim as linked or unlinked | Higher verified coverage wins |
| Editor verification | Claims awaiting re-check | Claims awaiting re-check | Publish only after every Xero billing claim is verified again |

For the first row, write the count of tasks completed correctly on the first attempt and the total tasks attempted, separately for each draft. Convert both pairs to rates before comparing them, because a bare success count against a different denominator is not a result. Apply the rule mechanically: the higher rate wins the row, and a row win here outweighs any impression formed while reading the drafts side by side.

For correction events, count each wrong, missing, or reversed action a reader had to repair — a field entered in the wrong place, a step skipped, an order of operations run backward. This row matters most when the success rates match: if both drafts pass the same number of tasks, the draft that required fewer repairs wins. A draft that arrives at a correct end state only after the reader undoes an earlier step has not earned the row.

Source traceability is a confirmation row, not the primary selector. Mark every Xero billing claim in each draft as linked to a source or unlinked, then compare verified coverage. Tool supply does not settle this: There's An AI For That lists 36 document comprehension AIs, so more drafting capacity is easy to obtain, while a claim attached to a checkable Xero billing source is not. Record coverage per draft rather than describing it.

Finally, treat the matrix as a selection, not an authorization. An editor re-verifies every Xero billing claim in the winning draft before publication, because a model's general reading score says nothing about your specific billing task — in a Washington Post review, five AI bots took a tough reading test and the strongest performer was not ChatGPT. The matrix picks the draft; the editor clears it.

## Count validation cost, not drafting speed

Speed is the cheapest number to collect and the most misleading one to rank on. For a Xero billing draft, review effort belongs inside the decision rather than in a footnote: a draft that appears fast and then absorbs verification, testing, and correction costs more than a slower draft that passes on the first attempt. Only a four-part time record shows that.

Log four values per draft, in the same unit, against the same frozen version: drafting or generation minutes, factual verification minutes, task-test minutes, and correction minutes. Total minutes per publishable draft = drafting + verification + test + correction. If a draft is regenerated, restart the clocks and attach the earlier minutes to the discarded version; carrying them forward flatters whichever draft survived. Compare publishable-to-publishable, never a verified draft against a raw one, or the totals measure different amounts of work.

| Time value | What it covers | Keep separate from |
| --- | --- | --- |
| Drafting or generation | Minutes to produce the draft | Verification and test minutes |
| Factual verification | Minutes to confirm each Xero billing claim against source material | Unsupported-claim count |
| Task test | Minutes to run the draft against the task it must enable | Correction minutes |
| Correction | Minutes to fix errors found by verification or testing | Drafting minutes |

Keep the error counts separate instead of collapsing them into one total: unsupported claims, wrong field names, missing prerequisites, and failed user actions each get a tally, and each correction event records the minutes it consumed. Separating them shows which category drives the total; a merged count hides it. In Xero billing material, a field name that does not match the product's label and a prerequisite the reader needed earlier can consume very different amounts of rework, and only separate counts make that visible.

Cost either error the same way. An unsupported claim must be re-verified against the source material, so its price is verification minutes plus correction minutes plus the test minutes for any task that ran on the disputed claim. A wrong field name or missing prerequisite must be fixed and the affected task re-tested, so its price is correction minutes plus re-test minutes. Both totals come out of review effort, not drafting, which is why drafting speed cannot pick the winner on its own. The Washington Post's review of five AI bots on its reading test found that the strongest performer was not the expected one — a reminder that fluency and accuracy are separate properties.

Totals fall only when the counts fall. Track each draft's total across successive revisions and check whether correction and verification minutes are dropping; if drafting minutes drop while the other three rise, the draft is not getting cheaper. Publish only after an editor re-verifies every Xero billing claim, and log that pass as its own entry rather than folding it into earlier numbers.

## Worked Example: Run the Numbers

Take one illustration and run it end to end. Scenario: a bookkeeper at a small accounting practice needs to answer a single Xero billing question — how to apply a credit note against an outstanding invoice — using nothing but the supplied draft. Party: one editor and one reader per draft, each working alone. Dates: a two-week window in March 2026. The counts below are stipulated so the arithmetic is visible; they are an illustration, not a reported study result.

**Step 1 — first-pass task success.** The editor-tested draft produced 10 attempts and 8 correct completions on the first pass: 8 ÷ 10 = 0.80, or 80%. The AI-generated draft produced 10 attempts and 6 correct first-pass completions: 6 ÷ 10 = 0.60, or 60%. Illustration A leaves the editor-tested draft ahead by 20 percentage points.

****Step 2 — correction events.** Count each wrong, missing, or reversed reader action that required repair. The editor-tested draft recorded 3 events across 10 attempts: 3 ÷ 10 = 0.30 per attempt. The AI-generated draft recorded 9 events: 9 ÷ 10 = 0.90 per attempt. Illustration B shows a gap of 6 events in total, or 0.60 per attempt.

| Metric | Editor-tested draft | AI-generated draft | Arithmetic |
| --- | --- | --- | --- |
| Attempts | 10 | 10 | equal base |
| First-pass successes | 8 | 6 | 8 − 6 = 2 |
| First-pass rate | 80% | 60% | 8/10 vs. 6/10 |
| Correction events | 3 | 9 | 9 − 3 = 6 |
| Corrections per attempt | 0.30 | 0.90 | 3/10 vs. 9/10 |
| Drafting minutes (illustration only) | 95 | 20 | excluded from selection |

**Step 3 — declare the winner.** The editor-tested draft wins this example: 80% beats 60% on the primary measure, and the correction count agrees (3 versus 9). A cross-check on the same inputs — first-pass successes minus correction events — gives 8 − 3 = 5 for the editor-tested draft and 6 − 9 = −3 for the AI-generated draft, so the two measures point the same direction rather than splitting.

**Break-even trigger.** The AI-generated draft reaches a tie at 8 ÷ 10 = 80%, which is one additional first-pass success. At that tie, the decision falls to correction events, so it must drop below 3 to win; at exactly 3 the tie persists. To win outright on the primary measure it needs 9 ÷ 10 = 90%. The 20 drafting minutes in the table do not move any of those thresholds. Before publishing whichever draft clears them, have an editor verify every Xero billing claim again.

## Run a six-task Xero billing trial

The trial starts with inputs you can hold constant: two drafts covering the same Xero billing workflow, one identical source packet feeding both, and one documented Xero version and date pinned at the top of every sheet. Fix six tasks before the first read-through and freeze them for the whole run. Each task gets three checkpoints, in order: locate the instruction, perform or specify the action, and explain the expected result. The order matters, because a draft can name the right screen and still fail when the reader never finds the field that opens it.

Record six-task results with the denominator visible. A draft that passes the locate checkpoint on four tasks is written as 4/6, not as a percentage, because the denominator is the entire trial. Seconds are logged per checkpoint as an observation and are not folded into the pass/fail score. Count a correction event only when the reader makes a wrong, missing, or reversed action that requires repair; describe the event separately from the count.

Use one worksheet per draft, per task:

| Field | Entry |
| --- | --- |
| Draft | ___ |
| Xero version/date | ___ |
| Task | ___ |
| Locate | pass/fail, seconds: ___ |
| Action | pass/fail, seconds: ___ |
| Expected result explained | pass/fail |
| Correction events | count: ___ |
| Notes | ___ |

A task counts as a first-pass success only when all three checkpoints pass on the reader's first attempt, with no help from the person running the trial. That single definition keeps the counts honest, because a task recovered on the second try is not a first-pass success even if the reader eventually lands the action. Drafting speed gets no column on the sheet.

Publication is conditional on the trial, not settled by it. The counts feed the comparison rule: higher first-pass task success decides, and a tie drops to the draft with fewer correction events. Before anything goes out, an editor re-verifies every Xero billing claim in the winning draft against the pinned version and date, including screen names, field labels, and each stated expected result. If a checkpoint shows the action cannot be performed in that version, treat it as a source-packet problem rather than a drafting score.

Six tasks in one version is a narrow instrument, and saying so is part of the method. The sheet describes what happened across those six tasks on that date; it says nothing about the Xero billing tasks you did not test. If the counts tie and the correction events tie, the trial has produced no winner, so return to the source packet and re-cut the tasks instead of publishing the draft that merely reads better.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | Put the AI-generated Xero billing draft and the editor-tested Xero billing draft through the same controlled task set from the 2026 comparison, and record each draft's first-pass task-success rate. | Identical tasks are the only way the two first-pass task-success rates can be compared without the drafting method skewing the score. |
| 2 | Compare the two first-pass task-success rates and mark the higher one as the winning Xero billing draft. | First-pass task success is the primary decision rule; a draft that finishes the task without help is the clearer one. |
| 3 | If the two first-pass task-success rates tie, count the correction events logged against each draft and select the draft with fewer correction events. | Correction events are the designated tie-breaker, so a tie is never resolved by preference or impression. |
| 4 | Set drafting speed aside entirely when ranking the AI-generated and editor-tested Xero billing drafts. | Speed is not part of the decision rule and would override the metric the guide is built on. |
| 5 | Hand the selected draft to an editor and have every Xero billing claim rechecked against current Xero product or API documentation, line by line. | Selection alone does not make a claim true; verification against current documentation is a publication gate. |

## Frequently Asked Questions

**What must a section do before it can be scored?**

It must map to an observable reader action, such as locating a billing setting, completing a step, interpreting an error response, or confirming that the account state changed.

**What happens if the reader must open another page to identify the action?**

Under the single-page rule, the section has not finished the sequence.

**What order should I use when comparing the drafts?**

Hold the order fixed as question, packet, draft, action, verification.

**Does producing an AI draft quickly move the comparison forward?**

Nothing advances because a draft was produced quickly.

**Who verifies the Xero billing claims before publication?**

An editor re-verifies every Xero billing claim against the manifest before publish.

**What did the Washington Post test show about relying on a model’s reputation?**

After reviewing five AI bots on one tough reading test, it found that the strongest performer was not the best-known model.

## Quick answers

| What should be done before comparing anything? | Convert each section into an observable action. |
| --- | --- |
| What question should be asked about each section? | Ask what the reader does after reading it. |
| What happens if a section maps to no action? | A section that maps to no action cannot be scored. |
| When has a section not finished the sequence? | If the reader must open another page to name the action, the section has not finished the sequence. |
| How does the pipeline end? | The pipeline ends when an editor re-verifies every Xero billing claim against the manifest before publish. |

Also worth reading: **Building documentation that actually helps your users**: [Building documentation that actually helps](https://specswriter.com/blog/building-documentation-that-actually-helps-your-users.php) · **The strategic reality of AI in remote technical documentation**: [strategic reality of AI in](https://specswriter.com/blog/the_strategic_reality_of_ai_in_remote_technical_documentatio.php) · **Skipping stakeholder review the riskiest shortcut in documentation**: [Skipping stakeholder review the riskiest](https://specswriter.com/blog/skipping-stakeholder-review-the-riskiest-shortcut-in-documentation.php)

### Related reading

- [Machine-Generated Documentation: 31% Task-Time Gain—Use Structure, Keep Controls](https://specswriter.com/blog/machine-generated-documentation-31-task-time-gainuse-structure-keep-controls.php)
- [How AI-Generated Websites Are Rewriting Technical Documentation Standards](https://specswriter.com/blog/how_ai_generated_websites_are_rewriting_technical_documentation_standards.php)
- [The Dark Side of AI Progress How OpenAI's O1 Model Raises New Questions About AI-Generated Portrait Authenticity](https://specswriter.com/blog/the_dark_side_of_ai_progress_how_openai_s_o1_model_raises_ne.php)
- [7 Essential Steps for Upscaling Midjourney-Generated Cyberpunk Animations to 4K Resolution](https://specswriter.com/blog/7_essential_steps_for_upscaling_midjourney_generated_cyberpu.php)
- [Australian Government Study Reveals AI Summaries Score 93% Lower Than Human-Generated Content in Document Analysis](https://specswriter.com/blog/australian_government_study_reveals_ai_summaries_score_93_l.php)
- [Nvidia's AI Chip Dominance 46% of Q2 Revenue Generated by Just 4 Customers in 2024](https://specswriter.com/blog/nvidia_s_ai_chip_dominance_46_of_q2_revenue_generated_by_ju.php)

### Latest

- [Machine-Generated Documentation: 31% Task-Time Gain—Use Structure, Keep Controls](https://specswriter.com/blog/machine-generated-documentation-31-task-time-gainuse-structure-keep-controls.php)
- [Duplicate Catalog Name: 1 Display-Name Match, No Approved Match, Reassign or...](https://specswriter.com/blog/duplicate-catalog-name-1-display-name-match-no-approved-match-reassign-or-remove.php)
- [Writing app documentation: Darwin Information Typing Architecture (DITA) vs bot...](https://specswriter.com/blog/writing-app-documentation-darwin-information-typing-architecture-dita-vs-bot-31-gap.php)

Canonical: https://specswriter.com/blog/xero-billing-guide-how-to-compare-ai-generated-and-editor-tested-drafts.php
Markdown: https://specswriter.com/blog/xero-billing-guide-how-to-compare-ai-generated-and-editor-tested-drafts.php/index.md
