Metadata Quality Assurance & Record Scoring
A catalogue’s value is bounded by the quality of the records in it, and quality in this context is not an abstraction: a record without an extent cannot be found by a map search, a record without a working distribution link cannot be used, and a record whose date means something different from its neighbours makes sorting meaningless. This topic sits inside the Metadata Catalog Automation & Ingestion Workflows framework and turns those observations into something measurable, reportable, and occasionally enforceable.
The failure it exists to prevent is a catalogue that is technically well-formed and practically unusable — every record valid against its schema, and a third of them impossible to find, evaluate, or download.
Score What Users Need, Not What the Schema Requires
Schema validity and usefulness are different properties, and a scoring model built from the schema’s mandatory elements measures the wrong thing. The useful model starts from what somebody does with a catalogue record and works backwards.
Weight the groups by consequence rather than equally. A record that cannot be found scores worse than one that can be found and lacks a lineage statement, because the first is invisible and the second is merely incomplete. In practice a weighting of roughly half for findability, a quarter for usability, and the remainder split between judgement and trust produces a score that ranks records the way a curator would.
A Score Is a Reporting Tool; a Gate Is a Blocking Tool
Conflating the two is the most common way a quality programme becomes unpopular. A score of 62 out of 100 is a useful thing to report and a terrible thing to block on, because nobody agrees where the line should be and the number moves whenever the model is refined. A gate should test specific, defensible conditions instead.
Keeping the model refinable matters more than it sounds. A score that blocks cannot be improved without renegotiating with every publisher whose records sit near the threshold, so it freezes — and a frozen quality model stops reflecting what users actually need within a year.
Report Per Source, Because That Is Where the Fix Is
An aggregate figure for the whole catalogue is almost useless operationally. Nobody can act on “the catalogue scores 71”, whereas “this department’s records score 38 and every one is missing a spatial extent” is a conversation with an obvious next step.
Group the report by source, then by rule, and order it by the number of records a single fix would improve. That ordering is what turns a quality report into a work plan: a mapping correction at the harvester that adds an extent for one source can lift thousands of records at once, and it is invisible in a report grouped by record.
Send it to the people who can act. A publisher who receives a quarterly summary of their own records — score, trend, and the three most common gaps — will usually fix the top item, because it is specific and it is theirs. The same information aggregated across the catalogue and sent to a platform team produces a backlog item nobody outside the team can complete.
Scoring Harvested Records You Do Not Control
For an aggregating catalogue, most records arrive from elsewhere and cannot be edited without breaking the next harvest. That constrains what quality work is possible, and it does not make measurement pointless.
Three responses are available and they should be used in this order. First, improve the mapping: a large share of apparent quality problems in harvested records are crosswalk artefacts — a field that exists upstream and is dropped, a date mapped from the wrong element, an extent that was expressed in a way the mapping did not recognise. Fixing the crosswalk improves every record from that source at once and costs nothing politically.
Second, enrich derivably: an extent can often be computed from the referenced service’s capabilities document, a format inferred from the distribution URL, a coordinate reference system read from the data itself. Enrichment must be recorded as such — a derived field marked as derived, so nobody later mistakes it for the publisher’s assertion.
Third, and only then, ask the source. A request accompanied by a specific list — these forty records lack an extent, here they are — is acted on far more often than a general observation about quality, and it is a request the provenance data described in the parent topic makes it possible to send at all.
Quality Decays Without Anybody Editing Anything
A record’s usefulness is not a property fixed at publication. It degrades on its own, through changes outside the catalogue entirely, and a quality programme that measures only at ingestion will report a stable, gradually more inaccurate picture.
Rescan the whole catalogue weekly rather than trusting the ingestion-time result, and report decay separately from ingestion quality — they have different owners and different remedies. A record that arrived poorly documented is a publisher conversation; a record that has decayed is usually a change somebody else made, and the catalogue’s job is to notice it before a user does.
Making the Programme Survive Its First Year
Quality initiatives are easy to start and unusually prone to quiet abandonment. Three things determine whether this one is still running in a year.
The first is that somebody owns the report. Not the tooling — the report, as a recurring piece of work with a name against it, whose job is to read it, decide what the top item is, and pursue it. A report generated automatically and read by nobody is indistinguishable from no report.
The second is that improvements are visible. Publish the trend per source, including the sources that improved, and mention them. The single most effective driver of metadata quality in a federated catalogue is a table where organisations can see each other’s scores, and it works far better as recognition than as pressure.
The third is that the model is allowed to change. What users need shifts — a new spatial search feature makes extents matter more, an open-data policy makes licences matter more — and a scoring model frozen at its first version gradually measures something nobody cares about. Version it, record when it changed, and treat a step in the trend as expected rather than as an anomaly to explain away.
Where Quality Rules Belong in the Pipeline
The same rule behaves very differently depending on where it runs, and placing rules deliberately is what keeps the pipeline from becoming either permissive or brittle. There are four useful positions and each suits a different kind of check.
At the source, before anything is transmitted, belongs whatever the publisher’s own editor can enforce: mandatory fields, controlled-vocabulary selection, format constraints. This is by far the cheapest place to fix a problem, because the person who knows the answer is present and the record has not yet been copied anywhere. It is also the position with the least reach for an aggregating catalogue, since the editor belongs to somebody else.
At ingestion, immediately after the crosswalk, belongs everything that decides whether a record can be stored coherently — an identifier, a parseable date, a valid geometry for the extent. These are the blocking rules, and running them here means a malformed record never reaches the catalogue rather than being cleaned up afterwards.
At indexing, when the record is prepared for search, belongs anything that affects discoverability: whether the extent is usable for a spatial query, whether keywords resolve against a vocabulary, whether the title is distinguishable from its neighbours. Failures here should not reject the record, because it is still valid; they should mark it as poorly discoverable and feed the score.
On a schedule, over the whole catalogue, belongs everything that can change without the record changing — the decay described above. These checks are the only ones that see a link that rotted last week, and they are the ones most often missing entirely.
Placing a rule at the wrong position produces a characteristic failure. A discoverability rule at ingestion rejects records that are legitimate but thin; a structural rule left to the schedule means malformed records sat in the catalogue for a week; and a decay check run only at ingestion reports a catalogue that was healthy at some point in the past.
Explaining a Score Somebody Disagrees With
Any scoring system will eventually produce a number a publisher disputes, and how that conversation goes depends entirely on whether the score can be decomposed. A single figure invites an argument about the model; a breakdown turns it into a list of specific, checkable statements.
Store the per-rule outcome for every record rather than only the total, and make the breakdown available alongside the score. “This record scores 44: no spatial extent, abstract under 40 words, no lineage, licence unspecified” is a sentence a custodian can act on or contest point by point, and either outcome is productive. If they contest it — the licence is stated in a field the crosswalk does not read, say — the finding is a mapping bug, which is the most valuable kind of feedback the programme can receive.
Two practices make this smoother. Keep the rule descriptions in plain language rather than in schema terms, because the audience is a data custodian rather than an engineer: “no area is given, so the record cannot be found by a map search” lands better than a reference to an element path. And show the rule’s weight, so it is clear why an absent extent costs more than an absent lineage statement — the weighting encodes a judgement about users, and exposing it invites exactly the discussion that improves it.
Finally, record disagreements and their resolutions. A publisher who successfully argues that a rule does not apply to their kind of dataset has identified a real gap in the model, and the next version should reflect it. A programme that never changes in response to feedback stops receiving any.
Starting Without Boiling the Ocean
A catalogue with thousands of records and no quality measurement can be improved substantially in a fortnight, and the way to do it is to resist building the full model first.
Begin by measuring one thing: the proportion of records with a usable spatial extent. It is the single highest-value field in a geospatial catalogue, it is trivially checkable, and the result is almost always worse than anyone expects. Report it per source, fix the crosswalk artefacts that the report exposes, and contact the two or three sources responsible for the rest. That alone typically moves a meaningful share of the catalogue from invisible to findable.
Then add the distribution-link check, because it is the second-highest-value field and the one that decays fastest. Then the abstract and keyword measures, which are softer and need a curator’s judgement to interpret. By the time all four are in place there is enough of a model to be worth calling a score, and — more importantly — there is a report with a history, which is what makes the trend meaningful.
Building the complete weighted model first is the alternative, and it usually consumes the enthusiasm the programme started with. The model takes weeks, it is disputed on arrival because nobody has seen intermediate results, and the first report lands as a large number nobody trusts. Incremental measurement produces smaller findings sooner, and each one is a conversation that makes the next step easier.
Operational Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Scores are stable and nothing improves | The report goes to a platform team rather than to publishers | Send per-source reports to the named custodian, with the top three gaps |
| A quality gate blocks a large harvest entirely | A rule promoted to blocking that is common upstream | Move it back to the score; blocking rules must be rare and defensible |
| Records score badly for fields that exist upstream | A crosswalk artefact, not a publisher problem | Fix the mapping first; measure again before contacting anyone |
| The score changed sharply with no content change | The model was refined without versioning | Version the model, and report a break in the series rather than a trend |
| Enriched fields are indistinguishable from asserted ones | Derivation not recorded | Mark derived fields; never let enrichment masquerade as an assertion |
| Everything passes and users still cannot find data | Findability weighted equally with completeness | Reweight from user tasks; missing extents outrank missing lineage |
| The report is too long to read | Grouped by record rather than by source and rule | Group by fix, ordered by how many records each one would improve |
FAQ
Should the score be visible to users of the catalogue?
A simplified form of it, yes — an indication that a record is thinly documented is useful to somebody deciding whether to rely on it, and it creates a mild incentive to improve. A raw number out of a hundred invites arguments about the model rather than improvements to the records; a small set of badges naming what is missing is more actionable and harder to dispute.
How many blocking rules should there be?
Few enough to list from memory: typically three to five. Every blocking rule is a promise to a publisher that their record will be rejected, and each one needs to be defensible in a conversation with somebody whose data is being refused. If the list is growing, that is usually a sign that the score is being used to argue for enforcement it should not carry.
What about records that are deliberately sparse?
Some are legitimately thin — a small reference dataset with a two-line description may be entirely adequate. This is why the score reports rather than blocks, and why the per-source view matters: a source with uniformly short abstracts for simple datasets is a different situation from one whose abstracts are empty, and only a human looking at the report can tell them apart.
Does an automated check replace curation?
No. It finds the mechanical problems — absent fields, broken links, malformed dates — reliably and cheaply, which frees curation to do the part no check can: judging whether an abstract actually describes the data, whether keywords are the ones somebody would search for, and whether the lineage statement would satisfy a person deciding to rely on it.
How often should scoring run?
On ingestion for the blocking rules, and on a schedule for the score, because a record’s quality changes without the record changing — a distribution link rots, a referenced service is withdrawn, a vocabulary term is deprecated. Weekly rescoring of the whole catalogue is inexpensive and catches the decay that ingestion-time checks cannot see.
Related
- Scoring Metadata Completeness in a CI Gate — the model and the gate, implemented.
- Deduplicating Harvested Records by Fingerprint — the quality problem aggregation creates.
- Detecting Broken Distribution Links in Catalog Records — the decay that happens after publication.
- CSW Catalog Schema Mapping & Validation — where crosswalk artefacts come from.
Up one level: Metadata Catalog Automation & Ingestion Workflows.