# Machine-Generated Documentation: 31% Task-Time Gain—Use Structure, Keep Controls

Brady Weaver · October 2, 2026

> A skeptical analysis of the claimed 31% task-time gain for machine-generated documentation, and why validation, controls, and evidence matter.

| Takeaway | Detail |
| --- | --- |
| 31% is a claim, not a benchmark. | None of the fetched sources reports a 31% result or attributes it to a named study, and no source compares machine-generated documentation, templates, and plain prose in the same experiment. |
| Drafting speed can hide validation work. | If the reported 31% task-time gain is real, validation may consume some or all of that difference before a developer trusts the API page. |
| Structure alone cannot establish superiority. | The 31% figure is not tied to comparisons among machine-generated, template-based, and unstructured prose. MariaDB’s hierarchy and search illustrate useful structure, but they do not establish a causal time or quality advantage. |
| Trust requires quality controls. | The reported 31% does not establish source fidelity, correctness, completeness, readability, comprehension, review cost, or maintainability; each outcome requires a separate measure. |

The reported 31% gain can create a tidy headline, but the supplied evidence provides no absolute task duration or baseline. Without that information, it cannot establish the difference between drafting and trusting the result. Does validation consume the claimed advantage before a developer can confidently use the API page? Speed matters, but a faster draft is not yet a correct, maintainable specification.

The supplied evidence does not support treating 31% as an established result. No fetched source reports that figure, names the study behind it, or compares machine-generated documentation, templates, and plain prose under one experiment. It also leaves “task time” undefined: drafting, editing, searching, reviewing, reading, understanding, and validating are different clocks. Without that boundary, the number cannot tell us what was actually faster.

Structure should still shape the workflow, but controls should determine release. Generated text can organize source material; templates can expose repeated fields; human review can check fidelity, edge cases, and maintainability. The winner is not whichever method posts 31%; it is the approach that produces documentation developers can find, understand, validate, and update without hidden labor. Keep source-grounded generation, explicit review criteria, and versioned examples—then measure the whole task, not only prose production.

![Machine-Generated Documentation](https://static.mm-ais.com/article-images-ai/machine-generated-documentation-31-task-ai-1eaefc28.jpg)

## The Time Equation

A fast first draft is not a delivered time saving. I separate drafting speed from delivered value with T_total = T_source + T_generate + T_validate + T_review + T_publish; I claim a time gain only when both arms incur all five stages and produce documentation with equivalent completeness. The supplied source set does not define whether “task time” means drafting, editing, reviewing, searching, reading, understanding, or validating documentation. According to the article headline in the supplied search-result set, the reported task-time gain over static templates is 31%; it is the only numerical performance claim, not a matched result. No fetched live-documentation source supplies a matched plain-text comparator or task-time report.

Free-form generation fails at a different objective. A language model predicts the next token most plausible from its context; it does not optimize for source-field completeness. Plausibility can therefore mask an omitted required parameter, a response definition contradicted elsewhere, or an invented example that reads fluently. The model’s objective gives it no reliable signal that its sentence violated an external source. Free-form output belongs in private, non-normative ideation; every published or execution-affecting document instead requires a version-controlled template with explicit source ownership.

Structured generation changes the objective. A source renderer owns machine-checkable fields, the model fills bounded explanatory slots, and validation checks the entire information architecture before publication. A prose block cannot count as complete merely because it sounds complete: every required block must be present, and every generated explanation remains tied to its owning source field. This is where faster prose should yield time rather than move ambiguity downstream.

JSON Schema Draft 2020-12 makes that boundary executable. Its required keyword turns an omitted field into a machine-detectable failure; type rejects a value of the wrong data type; enum rejects a value outside the allowed set. These are operational failures, not subjective prose criticisms. Schema success does not certify factual truth, so factual review remains a separate stage. The model can polish prose inside a valid slot, but it cannot satisfy a missing block by generating confident language.

| Stage | Matched timing boundary | Required record | Audit check |
| --- | --- | --- | --- |
| Source | Source snapshot to field inventory | Source map | Unowned published facts |
| Generate | Slot assignment to completed draft | Payload and slot log | Model work outside its slots |
| Validate | Invocation to pass/fail result | Validator report | Omission, type, or illegal-value failures |
| Review | Candidate opening to disposition | Defect and correction log | Schema and consistency defects |
| Publish | Release preparation to approved release | Versioned template and release record | Loss of source ownership |

The saving is credible only if the work ledger shows what disappeared: retrieval, boilerplate assembly, formatting, and first-draft composition. For each matched task, record those activities separately, then inspect whether they reappear during factual review, schema validation, or later maintenance. Maintenance is an edge case beyond publication: current speed cannot absorb deferred repair cost. Count an omitted parameter as a schema defect and a contradicted response or unsupported example as a consistency defect on both sides. The task-matched test should retain the reported task-time advantage over static templates while producing fewer schema and consistency defects than free-form prose. If it does, template-constrained generation wins; that joint result—not the obsolete-template myth that faster prose removes structure—makes structure the default for shippable technical documentation.

![The Time Equation — Machine-Generated Documentation](https://static.mm-ais.com/article-images-ai/machine-generated-documentation-31-task-ai-8c03aa94.jpg)

## Evidence Audit

The evidence supports acceleration, not an exemption from documentation controls. It justifies a task-matched test of template-constrained generation; it does not establish that free-form AI prose is reliable enough for API reference pages. Until that test isolates drafting speed from schema and consistency defects, faster output remains a hypothesis rather than a publication standard.

Before using the advertised task-time gain, require a provenance line naming the primary authors and year, sample size, comparator, measurement boundaries, and uncertainty. Do not silently reconstruct missing fields. If any element is unavailable, label the percentage “illustrative, not empirical.” This distinction prevents a result from a bounded writing or coding experiment from acquiring broader authority through repetition.

| Source | Verified result | Valid inference | Audit decision |
| --- | --- | --- | --- |
| Noy and Zhang’s 2023 Science experiment | The supplied evidence does not verify this study’s participant count, completion-time effect, or quality-rating effect. | No quantitative professional-writing effect is established by the supplied evidence. | Do not transfer the unverified result to API documentation. |
| Peng and colleagues’ 2023 GitHub Copilot experiment | According to Peng and colleagues, 95 professional developers completed a controlled JavaScript HTTP-server task 55.8% faster with Copilot. | AI assistance can accelerate a bounded software implementation task. | Do not transfer the result directly to reference-page production. |
| Stack Overflow’s 2024 Developer Survey | The supplied evidence does not verify this survey’s adoption or output-accuracy trust percentages. | No quantitative adoption-versus-confidence inference is established by the supplied evidence. | Require source validation; do not equate tool use with authority to publish autonomously. |

The inferential boundary matters. Even if Noy and Zhang identified an effect within their assigned writing tasks, those tasks would not establish conformity to a maintained documentation schema. The missing comparison is one in which required parameters, example validity, and cross-page status consistency are scored against an identified source owner.

Peng and colleagues move closer to software engineering, yet their HTTP-server task measures implementation completion rather than reference-page production. Faster code does not demonstrate faster or safer documentation. Stack Overflow’s survey supplies a separate warning: willingness to use AI cannot substitute for confidence in its output. The relevant control is traceability—who owns each source, validates each generated field, and records that decision.

The status-quo myth worth retiring is that faster prose has made templates obsolete. Missing structure is precisely what conceals omitted parameters, invented examples, and stale status codes. A version-controlled template turns those failures into visible schema obligations and review checkpoints; free-form output can appear complete while remaining unverifiable.

The concrete next action is to add the provenance line to the evidence record, then run the same documentation task under template-constrained generation and free-form AI prose. Keep task boundaries and defect scoring identical, and require blinded assessment of elapsed task time, schema validity, and consistency. The thesis advances only if the template condition preserves the advertised time advantage while producing fewer defects. For every published or execution-affecting document, use a version-controlled template with explicit source ownership. Reserve plain machine prose for private, non-normative ideation.

![Evidence Audit — Machine-Generated Documentation](https://static.mm-ais.com/article-images-pixabay/machine-generated-documentation-31-task-9faeeabc.jpg)

## Production Verdict

Ship by structure, not fluency. Template-constrained generation should be the production default for shippable technical documentation, conditional on a task-matched test showing that it retains the reported task-time gain documented above while producing fewer schema and consistency defects than free-form AI prose. The governing rule is strict: every published or execution-affecting document uses a version-controlled template with explicit source ownership; plain machine prose remains limited to private, non-normative ideation.

This is a falsifiable production verdict, not a style preference. No source in the supplied set compares all three modes in the same experiment. According to the supplied record for arXiv:1401.5353v1, its documentation corpus is not identified as handwritten, template-based, or machine-generated; its page count therefore cannot establish a method advantage. The missing evidence is a controlled comparison in which structure, rather than eloquence, determines the result.

Run a 2026 three-arm pilot across the same API operations and source snapshot: plain generation, static templates, and template-constrained generation. Randomize page order, and blind reviewers to condition. Give each reviewer the same task the page governs—for example, obtaining a required credential, sending the correct request, interpreting the response, or executing a migration step. Score whether the reader performs the correct action, not whether the text is fluent or merely recognizable as plausible.

Measure five outcomes on those same pages: time to publish; factual or schema defect rate across the reviewed fields; reviewer-minutes per page; percentage of source changes propagated automatically; and first-read task success. Hold content scope, source access, tooling, and reviewer expertise constant. Prose preference can remain a diagnostic, but it is not a quality measure and cannot override a failed task, an unowned field, or a schema violation. Break out results by field type and publication risk so an aggregate improvement cannot conceal a consequential omission.

The structured path should win only if it preserves the documented time advantage while improving correctness and action. Its review boundary is mechanical: source-owned values enter named slots, unsupported claims fail validation, and regeneration with a source diff exposes drift. Plain prose generates text quickly but makes omissions and inventions difficult to localize; static templates are inspectable but still depend on manual synchronization. This mechanism is why structure is the durable control and an LLM’s stylistic range is not.

Apply the verdict by publication risk. High-consequence API, specification, security, and migration pages require the template-constrained path after it passes the release test. Static templates remain acceptable for stable, deterministic one-offs when maintainers own each update. Unconstrained prose remains a prepublication option only for private, non-normative exploration. Operationally, the release gate is simple: if a page lacks a version-controlled template, a named source for every normative field, or a reviewable source diff, it is not ready to ship.

| Option | Draft speed | Source fidelity | Review burden | Maintenance | Verdict |
| --- | --- | --- | --- | --- | --- |
| Plain machine prose | Fast text generation | Variable; source fields may be omitted or invented | High and difficult to localize | Prose searches and manual comparison | Reject for shipped normative docs |
| Static template | Medium | High only after manual synchronization | Low for stable pages | Manual updates can drift | Viable for deterministic one-offs |
| Template-constrained generation | Fast | High for source-owned fields | Medium and bounded to named slots | Regeneration and source diffs | Winner for shippable technical documentation |

![Production Verdict — Machine-Generated Documentation](https://static.mm-ais.com/article-images-pixabay/machine-generated-documentation-31-task-148918aa.jpg)

## Counter-Evidence

The strongest counter-evidence is a scope mismatch: the cited empirical benchmarks do not test production API references or software specifications. I therefore treat their time effects as directional priors, not direct estimates of the headline gain. According to the article headline and supplied source set, no fetched source attributes that gain to a named study; the accessible ResearchGate records expose topics but no sample sizes, scores, effect sizes, or implementation details. The available documentation records also do not compare generated, templated, and plain prose. This is an evidentiary boundary, not evidence of equivalence.

A drafting benchmark becomes misleading when it stops at generated paragraphs. If it excludes source extraction, schema validation, factual review, CI checks, accessibility testing, localization, and maintenance, its outcome measures prose production rather than a complete documentation cycle. The omitted stages are precisely where template-constrained generation should expose omissions and inconsistencies. Calling that narrower measure a delivered-time result confuses output creation with publishable work.

Structure can also preserve wrongness. A static template can carry an obsolete endpoint, deprecated option, or stale status code from one release into the next. Required fields may all be present while the engineering source is already wrong; schema conformance establishes structural consistency, not source currency. The practical control is explicit source ownership plus publication-time validation against the authoritative specification or service contract, with a failed source check blocking release.

The format premium weakens for conceptual tutorials and architecture explanations. Forcing every paragraph into identical slots can interrupt transitions, hide uncertainty, and give readers a sequence of components rather than an argument. Narrative coherence is therefore a legitimate counterweight to uniform reference formatting. This does not justify free-form publication: the template can require sections, evidence, and accountable sources while allowing rhetorical variation inside them.

Aggregate means can conceal the users most likely to be harmed. Results may vary with domain expertise, language, accessibility needs, and familiarity with the API. I would report median and 90th-percentile task time alongside incorrect-action rates, disaggregated by those conditions. A fast median cannot compensate for a long upper tail or a subgroup repeatedly taking the wrong action; comprehension and accessibility are part of task performance, not decorative secondary metrics.

Reproducibility is the final brake. A result labeled only by vendor and model family is not reproducible in 2026. Preserve the exact model snapshot, system prompt, retrieval corpus, generation settings, validation tools, and evaluation date, because silent model updates can move the baseline. The ResearchGate page titled “Correlating Automated and Human Evaluation of Code Documentation Generation Quality” returned only a CAPTCHA, while Microsoft’s “Bringing the power of machine reading comprehension to specialized documents” exposed a sign-in or error shell rather than its methods or results. Neither inaccessible record establishes a performance result.

The thesis therefore remains conditional, not mystical: template-constrained generation should retain the reported task-time advantage while reducing schema and consistency defects in the relevant task. If it does not, the time claim fails for that case; it does not make prose the safer production default. Keep a version-controlled template and explicit source owner for every published or execution-affecting document, permit plain machine prose only for private, non-normative ideation, and revise rigid slots or stale sources before release.

![Counter-Evidence — Machine-Generated Documentation](https://static.mm-ais.com/article-images-pixabay/machine-generated-documentation-31-task-206f4d23.jpg)

## Worked API Draft

A fluent issue-creation guide is an unreliable production artifact because prose can hide a missing required field. According to GitHub REST API documentation’s “Create an issue” operation and “Versioning the REST API” guidance, build an OpenAPI 3.1.0 contract for POST /repos/{owner}/{repo}/issues. Pin X-GitHub-Api-Version to the dated compatibility value 2022-11-28, then revalidate that value against GitHub’s current versioning guidance before publication.

| Contract element | Source-locked record | Provenance test |
| --- | --- | --- |
| Operation | POST /repos/{owner}/{repo}/issues | Method and path match GitHub exactly |
| Version | OpenAPI 3.1.0; header value 2022-11-28 | Both values are explicit, not model-selected |
| Path parameters | 2 required strings: owner and repo | Neither parameter is omitted or renamed |
| Required request field | title, string, required | Requiredness comes from the operation |
| Success response | Source-defined success response | Created status matches the source |
| Failure responses | 404 plus other source-documented failures | Every documented response remains in the error matrix |

Freeze the rendered endpoint page into exactly eight blocks: purpose; authorization; path parameters; request fields; success response; error matrix; executable request example; and version or rate-limit links. This immutable set is the machine-documentation template under test—not a menu from which the generator may select convenient material.

Commit the contract and eight-block template together with an owner map. GitHub’s published operation remains canonical for parameter types, required fields, status codes, and headers. The model may draft only descriptions and example prose. Error interpretation and security behavior require engineer review. Plain machine prose remains private, non-normative ideation; it cannot become a published or execution-affecting document without passing through the owned template.

Only after the primary-source provenance check passes, compare the static-template and constrained workflows using the same timing boundary. The supplied material provides no absolute task duration or template baseline, so the reported percentage cannot be converted into minutes or hours. If the reported clock stopped at draft generation, separately disclose the two planning assumptions used in the worked example: 8 minutes for Spectral and schema validation plus 7 minutes for engineer and security review. Those additions are planning assumptions, not observed measurements. If the reported clock already included validation and review, do not add them again.

The production claim requires a same-task comparison using the same schema and consistency defect taxonomy. Template-constrained generation earns the time advantage only with zero schema errors and evidence that free-form prose produced more defects. Plain prose without equivalent defect data cannot be scored.

| Recorded condition | Constrained cycle | Static baseline | Decision | Reason |
| --- | --- | --- | --- | --- |
| Any primary-source mismatch | Not timed | Not timed | Stop | Repair provenance before comparison |
| Validation and review already included; schema errors equal zero | Claimed acceleration only | Not supplied | Not empirically scoreable | No absolute duration or matched study is supplied |
| Draft-only clock plus 8-minute validation and 7-minute review; equivalent free-form defect data show more defects | Not established | Not supplied | Not scoreable | Planning assumptions do not establish a net gain percentage |
| Post-generation review consumes the full claimed draft gap | Not established | Not supplied | Not scoreable | Review may erase the reported acceleration |
| Post-generation review exceeds the claimed draft gap | Not established | Not supplied | Not scoreable | Additional review may outweigh acceleration |
| Free-form arm lacks equivalent defect data | Not established | Not supplied | Not scoreable | Time alone cannot establish a defect advantage |

![Worked API Draft — Machine-Generated Documentation](https://static.mm-ais.com/article-images-pixabay/machine-generated-documentation-31-task-c563c23a.jpg)

## Choose by Ship State

Ship state, not prose velocity, sets the generation boundary. A fluent answer is not evidence that it is publishable. Free-form AI does not make templates obsolete; missing structure is what can conceal omitted parameters, invented examples, and stale status codes.

At the publication state, if readers will implement, configure, integrate, migrate, or support software, require a version-controlled template with explicit ownership by the team or system controlling the canonical source. Permit plain machine prose only for a private outline, private alternative-title candidates, or non-normative brainstorming. A published or execution-affecting document receives no naked-prompt exception.

Source ownership is field-level. If a value, field, endpoint, option, status, version, or code sample exists in a canonical machine-readable source, render it deterministically. Store generated language only in the bounded description slot attached to that field; the model must not rename, rewrite, infer, or repair the sourced artifact. A missing source key should block rendering rather than invite the model to fill the gap.

At the consequence state, if an error could affect authentication, authorization, billing, data loss, security, or a destructive migration, require an executable example test and sign-off by a named engineer before publication. Sampled review is insufficient. For example, an OAuth authorization-server guide needs an executed request path and a named engineering approver; visual fluency cannot substitute for either control.

Use change load to select the renderer. If a fact changes at least monthly or appears in three or more public surfaces, generate each occurrence from one canonical source. An OAuth status shown

## Frequently Asked Questions

**What five stages must be timed before a drafting-speed claim counts as a delivered time saving?**

A matched claim must time T_source + T_generate + T_validate + T_review + T_publish for both arms and require equivalent documentation completeness.

**Can the advertised 31% task-time gain be treated as an established empirical result?**

No, because the supplied source set neither attributes the 31% figure to a named study nor reports a matched comparison, and it leaves “task time” undefined.

**What provenance is required before using the advertised 31% task-time gain?**

It must name the primary authors and year, sample size, comparator, measurement boundaries, and uncertainty; if any element is unavailable, the percentage must be labeled “illustrative, not empirical.”

**How does JSON Schema Draft 2020-12 detect missing or invalid structured documentation content?**

The `required` keyword flags an omitted field, `type` rejects a wrong data type, and `enum` rejects a value outside the allowed set, although schema success does not certify factual truth.

**What must the task-matched test show before template-constrained generation is considered the winner for shippable API documentation?**

It must retain the reported 31% task-time advantage over static templates while producing fewer schema and consistency defects than free-form prose.

**Where may free-form model output be used, and what is required before publication?**

Free-form output is limited to private, non-normative ideation, while every published or execution-affecting document requires a version-controlled template with explicit source ownership.

## Quick answers

| Does the supplied evidence establish the reported 31% task-time gain? | The supplied evidence does not support treating 31% as an established result. |
| --- | --- |
| What should determine whether generated documentation is ready for release? | Structure should still shape the workflow, but controls should determine release. |
| Where does the article say free-form generated output belongs? | Free-form output belongs in private, non-normative ideation; every published or execution-affecting document instead requires a version-controlled template with explicit source ownership. |
| What failures can JSON Schema detect? | Its required keyword turns an omitted field into a machine-detectable failure; type rejects a value of the wrong data type; enum rejects a value outside the allowed set. |
| Does schema validation certify factual truth? | Schema success does not certify factual truth, so factual review remains a separate stage. |

Also worth reading: **The strategic reality of AI in remote technical documentation**: [strategic reality of AI in](https://specswriter.com/blog/the_strategic_reality_of_ai_in_remote_technical_documentatio.php) · **Skipping stakeholder review the riskiest shortcut in documentation**: [Skipping stakeholder review the riskiest](https://specswriter.com/blog/skipping-stakeholder-review-the-riskiest-shortcut-in-documentation.php) · **Building documentation that actually helps your users**: [Building documentation that actually helps](https://specswriter.com/blog/building-documentation-that-actually-helps-your-users.php)

### Related reading

- [NVD: Documentation Structure, Not Data, Boosts CVE Recall 38%](https://specswriter.com/blog/nvd-documentation-structure-not-data-boosts-cve-recall-38.php)
- [How AI-Generated Websites Are Rewriting Technical Documentation Standards](https://specswriter.com/blog/how_ai_generated_websites_are_rewriting_technical_documentation_standards.php)
- [Writing user guides 2026: Application Programming Interface (API) 32% machine cut vs keep](https://specswriter.com/blog/writing-user-guides-2026-application-programming-interface-api-32-machine-cut-vs-keep.php)
- [Writing app documentation: Darwin Information Typing Architecture (DITA) vs bot 31% gap](https://specswriter.com/blog/writing-app-documentation-darwin-information-typing-architecture-dita-vs-bot-31-gap.php)
- [Building Product Documentation and White Papers with AI Prompts](https://specswriter.com/blog/building_product_documentation_and_white_papers_with_ai_prompts.php)
- [Automate AI Documentation with GitHub and Azure DevOps](https://specswriter.com/blog/automate_ai_documentation_with_github_and_azure_devops.php)

### Latest

- [Duplicate Catalog Name: 1 Display-Name Match, No Approved Match, Reassign or...](https://specswriter.com/blog/duplicate-catalog-name-1-display-name-match-no-approved-match-reassign-or-remove.php)
- [Writing app documentation: Darwin Information Typing Architecture (DITA) vs bot...](https://specswriter.com/blog/writing-app-documentation-darwin-information-typing-architecture-dita-vs-bot-31-gap.php)
- [How to write docs 2026: Darwin Information Typing Architecture (DITA) vs...](https://specswriter.com/blog/how-to-write-docs-2026-darwin-information-typing-architecture-dita-vs-freeform-31-cut.php)

Canonical: https://specswriter.com/blog/machine-generated-documentation-31-task-time-gainuse-structure-keep-controls.php
Markdown: https://specswriter.com/blog/machine-generated-documentation-31-task-time-gainuse-structure-keep-controls.php/index.md
