Method

Scoring AI utility prose

May 14, 2024

Six checks I use before I trust a sentence about a service line.

I spent eight years checking service-line records for a gas utility. The work was not writing. It was deciding whether a feature on the map was allowed to stay. A field packet came back from the crew. The packet said what was in the ground. The GIS — GE Smallworld, and later checks I also ran in ArcGIS — said what the system believed. When those two disagreed, a status sentence did not get a vote. The packet did.

I use that habit on AI-generated utility prose. A model can produce a paragraph that sounds like a close-out comment. The verbs are confident. The nouns are the right family: service, as-built, asset, network. Nothing in the paragraph is obviously wild. That is the failure I care about. The draft is finished in tone and unfinished as a record.

I do not start by rewriting. I score. Six checks, in order. If a check fails, I mark it and leave the sentence alone until the set of marks is complete. Repair comes after the score, and the repair may touch only what a mark named.

Source of record

A utility sentence should let a second reader name the thing it depends on. That thing is a system or a document: the LocusView as-built, the field packet, the work order, the feature in Smallworld or ArcGIS. “The documentation has been updated” fails this check. Updated where, and from what? If I cannot point, the sentence is not yet a statement. It is a gesture toward completeness.

I mark that failure SOR. The repair is not a smoother verb. The repair names the object. “The LocusView as-built in the packet has not been posted to the service feature” can be checked. A person can open the packet and open the feature. The first sentence cannot.

Load-bearing words

In a service-line note, a few words carry the claim. Material. Status. Replaced or retired. Abandoned in place or removed. Proposed or installed. I underline those words and ask whether the source used them. Models prefer softer cousins: updated, aligned, reflected, captured. Those words survive a casual read because they do not commit. On a gas map they are risky for the same reason. They sound like work got done.

If the packet says the old steel was abandoned in place and capped at the main, the note has to be willing to say that. “The legacy asset was addressed” is a TERM failure when “abandoned in place” was available and was the value that changes the feature. I put the system word back. I do not upgrade it into a story about a program.

Fake precision

A station, a length, a diameter, or a date is either in the source or it is an invention. Drafts grow a number because a number makes a paragraph look like GIS. I mark that INV. In a training sample I either keep the figure that was given or I write “station as shown on the as-built” and stop. A made-up half-foot is worse than a short list of fields, because a new analyst will treat it as a standard.

Followability

Someone has to do the next thing. A mapper posts a feature or holds it. A coordinator tells a crew the packet is incomplete. A reviewer knows whether the order can close. If the paragraph does not name that next action, it is a summary of a mood. “Ensure data quality before close-out” is not followable. Compare material, station, and abandon status, and if they disagree, stop: that is followable.

I mark the unfollowable line STEP, even when it is one sentence rather than a list. A procedure hiding inside a paragraph still needs a stop. “As needed” is the stop, disguised as an adverb. The repair pulls it out and writes it as its own instruction: if the three items disagree, do not edit the feature.

Audience

The same facts do not belong in the same order for every reader. A mapper needs the attribute and the geometry cue. A coordinator needs the status and what is waiting. I ask who acts on this text. If the draft answers a person who is not in the room — often “stakeholders” — I mark AUD and split the text rather than average it.

Two short notes, each in the order that reader works, are a better repair than one averaged paragraph.

Leftover AI

Some sentences are not wrong about a fact. They are padding that signals “this is complete.” I keep a short list, and I quote it rather than imitate it: “in today’s rapidly evolving landscape,” “robust single source of truth,” “seamless integration,” “enhanced operational efficiency.” None of those phrases can be checked against a packet. They fail even when every fact around them is fine, because they teach a model that fluency stands in for a source. The tag is PAD. The repair is to cut them, not to translate them into a tasteful equivalent. There is no accurate version of a sentence that was only there to sound finished.

What the score is for

The score is a constraint on the rewrite. When I repair, I do not “make it sound more human.” I change the marked spans. A PAD sentence is cut. A missing source is named. A soft verb is replaced with the verb the record used. If I notice a nicer rhythm while I am in the sentence, I leave it unless a check required the change. Polish without a mark is how a second error gets in.

Show the bad paragraph and the marks, not only the clean result. The clean result hides the decision. The marks are the decision. A claim may remain when a named record supports it, and it is held when the record does not. Sounding finished is cheap. Being true is a check.

All writing