📋 PLAYBOOK

A dark wall of archival card catalogue drawers, one drawer open and lit from within by a blue glow, a blank paper tag hanging from its handle.

By the end of this playbook you will be able to run; configuration standards that remain as your estate grows, a smaller set of records that responders trust because every one of them has a named owner, a ruled source of truth and a recent human check behind it. The move is not to record more of the estate. It is to govern less data, harder.

When to use this. You have a configuration database in place, or are standing one up, and more than one team feeds it. You have at least one automated population source running, and leadership willing to put names against data. If you are still designing the data model itself, do that groundwork first, because this playbook assumes the model exists and asks how you keep what fills it trustworthy.

Stage 1. Put names against the data before you touch a record

Every configuration estate that stays healthy has four kinds of owner, even if at least one of them is missing, your estate will decay. Someone must own the standards themselves, enforcing the rules and settling disputes about whose data wins. Someone must own the structure, deciding which classes of item exist, which sources feed them, and how conflicts between those sources resolve. Someone must do the implementing, or the whole design stays on paper. And each domain needs a steward, so your server lead, network lead and database lead each own the attribute requirements and the data quality for their own patch, because they are the only people who can tell a real record from a plausible one.

Skip this stage and you’re accepting the following; Disputes about ownership stalling for months because nobody is empowered to settle them, classes multiply because nobody is empowered to say no, and the teams closest to the infrastructure quietly stop correcting records they were never asked to own.

The gotcha is assigning stewardship to the platform team because they run the tooling. The platform team can tell you whether a record parses, but only the domain team can tell you whether it is true.

Stage 2. Shrink the scope until completeness is achievable

A modern platform can model hundreds of distinct classes of item, well over seven hundred at the last count, and governing all of them is not thoroughness but a guarantee that none of them are governed. Start with the dozen or so classes that actually carry your incident and change traffic, which usually means servers, network equipment, databases and the datacentre estate, and let the rest wait until this core has earned trust.

Then tier the attributes within those classes by who needs them and when. The incident-time set is what a responder reaches for while something is down, meaning the item's name, its address on the network, its operating system, its support group and its owner. The planning set supports reporting and forward decisions, covering location, environment and business criticality. Everything beyond those two tiers is optional until proven otherwise, and any custom addition should carry a written business case, because every manually maintained field is a standing debt someone must keep paying.

⚠️ COMMON PITFALL

Marking fifty attributes as required feels rigorous and works like sabotage. Completeness never climbs above forty per cent, the score stops meaning anything, and teams stop caring. Cap the required set at five to seven attributes per class, get them to ninety five per cent, and only then add more.

Stage 3. Rule which source wins, attribute by attribute

Your records are fed by automated discovery scans, an asset register, cloud provider feeds and manual entry, and those sources will disagree. When they do, the platform resolves the conflict somehow, and if you have not written the rule, the resolution is arbitrary and nobody can explain it afterwards.

The workable rule is precedence by attribute type, not by source. Technical facts belong to automated discovery, because a scanner that has just inspected the machine knows its address and operating system better than any human record. Operational facts belong to the operations teams, because support group, criticality and ownership are organisational decisions no scanner can observe, and they must be protected from automated overwrite. Commercial facts, such as purchase dates and warranty terms, belong to the asset system that transacted them.

Attribute type

Examples

Authoritative source

Technical

Network address, operating system, installed software

Automated discovery

Operational

Support group, criticality, ownership

The operations teams, protected from overwrite

Commercial

Purchase date, warranty, asset reference

The asset register

Then publish the escalation ladder for the disagreements the rule does not settle. Stewards resolve attribute disputes inside their own domain, the standards owner arbitrates across teams, and the sponsor settles matters of policy. Decide each dispute once, write the outcome down, and move on.

The gotcha is the one-source policy, where automated discovery simply wins everything. It lasts until the first time a team watches its careful manual corrections get wiped by an overnight scan, after which they stop maintaining the data at all and the estate rots politely from the inside.

Stage 4. Set a measurable bar and check it monthly

Without a stated quality bar, you find out the data is bad when a responder pulls up a record mid-incident and the support group is wrong. Define good in advance, across three failure classes.

  • Records missing their required fields.

  • Records that haven’t been updated within a stated window, and that window should be measured in days matched to your scan cadence, not in months.

  • Duplicates, which deserve more fear than they get, because a five per cent duplicate rate across forty thousand records is two thousand places where a responder can pull up the wrong one.

Then verify with humans every month, with each steward pulling twenty records from their class and checking them against sources they know to be good, aiming for ninety five per cent correctness on the incident-time attributes. It is a thirty minute exercise, and it catches drift while it is still cheap to correct.

The gotcha here is dashboard theatre. A health score that is watched but never worked is decoration, so treat every failed check as a queue item routed to a named owner with a date, or stop pretending to measure.

Stage 5. Make someone vouch for every record on a cadence

Automated population tells you what the scanners saw, but it cannot tell you whether the thing should still exist, who really supports it, or whether it matters. So put a vouching cadence on the estate, a written policy stating which records must be confirmed by a person, how often, and by whom. Assign the confirmations to the people who manage the items rather than to whoever has spare capacity, and give rejection teeth, so that when a steward says this thing no longer exists, the record leaves the estate.

This mechanism is what converts your configuration data from what the system remembers into what a person recently vouched for. Spend the human attention where it pays, because records your scanners verify every night can be confirmed automatically, which frees your stewards for the records no scanner can see.

The gotcha is that vouching without knowledge is rubber-stamping. Assign confirmation tasks to people who cannot recognise the items, or drop hundreds on one person in a single batch, and you will get a hundred per cent approval rate that certifies nothing.

Stage 6. Expand in rings, and validate each ring before the next

Scope creep kills more configuration programmes than technology does, so expand in rings, in value order. Core infrastructure comes first, because that is where incident and change work sees immediate benefit, then applications and middleware, then the cloud estate, with end-user devices last and only if a process actually needs them.

At each ring, run the new sources against a small sample before scaling, asking whether they report what you expected and whether duplicates are appearing, and share what you find with the stewards who will own the results before you widen. A mistake caught at fifty records is an afternoon. The same mistake at fifty thousand is a remediation programme.

The gotcha is scoping everything on day one. The estate grows faster than the governance, and standards that were never enforced anywhere end up enforced nowhere.

The checklist

  1. Name the four owners (standards, structure, implementation, and a steward per domain).

  2. Pick the classes that carry incident and change traffic, and ignore the rest for now.

  3. Cap required attributes at five to seven per class, tiered by who needs them and when.

  4. Write source precedence per attribute type (technical to discovery, operational to the teams, commercial to the asset register).

  5. Publish the escalation ladder (steward, then standards owner, then sponsor).

  6. Set the staleness window in days, matched to scan cadence, and a ninety five per cent bar on incident-time attributes.

  7. Run the monthly spot check of twenty records per class against known-good sources.

  8. Put a vouching cadence on everything a scanner cannot see, with rejection that retires records.

  9. Expand one ring at a time, and sample-validate before each ring.

Common failure modes

Everything is required. Completeness collapses, the score loses meaning, and teams disengage. The fix is to cut to five to seven attributes, reach ninety five per cent, and then add.

One source rules everything. Manual corrections get overwritten, trust collapses, and maintenance stops. The fix is precedence by attribute type, with operational facts protected.

The estate only ever grows. Retired systems linger as records because nothing removes them. The fix is a vouching cadence where rejection actually deletes.

Governance by dashboard. The score is reported and nothing changes. The fix is turning every failed check into a routed task with a named owner and a date.

The bottom line

If you protect only one thing from this playbook, protect named ownership per class. Standards enforce themselves when a specific person answers for a specific slice of the data, and they enforce nothing when quality belongs to everyone in general.

A configuration record is an asset only while someone will vouch for it. The rest is archaeology.

P.S. Identify the one attribute your teams argue about most. The answer usually reveals which owner is missing.