Your product data is a mess. Which parts are worth fixing once?
Every supplier sends a different format, half the attributes are missing, the feed throws errors and nobody has ever agreed what the categories mean. That is the ordinary condition of an apparel catalogue and it is not a failure of discipline. What is worth knowing is that only two of your attributes pay off in more than one place, three of them will waste your time by looking as though they do, and one of them is permanent work that never finishes.
Navigate this page
The short answer
Sort your attributes by how many destinations one correct value serves. There are four groups and they behave completely differently.
| Group | Attributes | Fix it once? |
|---|---|---|
| Universal | The product identifier, and the party record about your company | Yes. One correct value serves every destination there is. |
| False friends | Material, weight, brand | No. They exist in every system under the same name and mean different things. |
| Single destination | Colour, size, gender and age group on one side; recycled content, care instructions and producer registration on the other | No, and that is fine. Fix them where they are used. |
| Never | Product category | It has no crosswalk. Maintaining one is a standing job, not a project. |
Most catalogue clean-up effort goes into the last three groups, because that is where the errors are visible. The gain is in the first.
The two that genuinely pay off once
The product identifier. It is the only universal join key in the whole chain. It is present in commerce feeds, in marketplace listings, in retailer onboarding, in the open web vocabulary, and it is what a passport hangs off. It is also the field merchants most commonly cannot supply, and the one where an error becomes permanent, because a code goes onto goods and the goods go into somebody's wardrobe. What allocating one commits you to, how many a range actually needs and what it costs by country is a subject of its own.
The party record. Who made this, who imports it, who answers for it inside the Union, and which producer schemes the company is registered with. It is small, it is stable, it changes rarely, and it applies to a whole catalogue rather than to a product. Every system that can stop you trading asks for it, and almost every system a merchant actually uses stores it in the wrong place, which is why it gets typed again every time. That is a separate subject and it is on the registrations that gate a listing.
Between them those two carry more reuse than every other attribute combined. They are also the two nobody enjoys working on, because neither of them makes a product page look better.
The three that waste your time by looking universal
These appear in commerce systems and in regulatory formats under the same name and describe different objects. They fail silently in both directions, which is what makes them expensive: a value good enough for a feed is not a compliance value, and a compliance value pushed back the other way arrives as a string the feed cannot use.
Material. A shopping feed wants a primary material as a short string, used to filter and display. Textile labelling wants a percentage per fibre, drawn from a closed statutory list, held per component so that a shell, a lining and a trim do not collapse into one word.
Weight. A commerce system means shipping weight, for rate calculation. Producer responsibility reporting means mass placed on the market, which is the basis of a fee.
Brand. A feed wants a marketing string. The regulatory side wants a legally accountable entity with an address.
The full account of why one word ends up meaning two things, and the three harder pairs sitting underneath these, is on where product data stops meaning the same thing. What matters for planning is simply that a clean-up project which maps by field name and calls that reuse will under-deliver on one side and over-claim on the other. Why one word ends up meaning two things is the subject of where product data stops meaning the same thing.
The ones that only ever serve one destination
Worth naming, because a lot of effort goes into treating them as though they should reconcile.
Commerce only. Colour, size, gender and age group appear in no regulatory instrument. Colour normalisation against an enumerated marketplace list is a real and irritating problem with no compliance counterpart at all. Size on some marketplaces is a structured hierarchy of gender, age range, size system, size class and body type, and none of that reaches a passport.
Regulatory only. Recycled content, recyclability and care instructions have no standard home in any commerce or search feed specification we checked. That data has to be sourced, verified and maintained entirely outside the commerce stack, and there is no reuse argument in either direction. Producer registration numbers are the same shape and worse, because they are not a product attribute at all.
The practical consequence is a relief rather than a problem. You do not have to reconcile these. You have to know which side of the line each one sits on, and stop paying twice to make them agree.
The one that never finishes
Product category classification is the starkest non-overlap in the whole picture. A global classification standard, a shopping taxonomy, a marketplace browse tree, a passport's own scheme and a different national schedule for every producer responsibility scheme are five or more mutually non-mappable hierarchies for the same underlying concept.
There is no mapping. Anything claiming to structure once and publish everywhere has to maintain a crosswalk per destination, and that crosswalk is permanent maintenance rather than a one-off. Budget it as a standing cost or do not promise the outcome.
What nobody can tell you, including us
The claim underneath this whole subject is that the same attribute gets collected repeatedly inside one company for different purposes. It is almost certainly true and no study anywhere measures it.
The EU-funded project that studied passport costs for smaller businesses directly reported that providers themselves would not supply quantitative figures, for commercial reasons and in some cases because they did not know. The widely quoted claim that one in four product records is wrong has no primary document behind it; the nearest real source is a UK grocery data quality report from 2009, measuring a different construct in a different sector. Two vendor figures for returns caused by poor content differ from each other and neither publishes a sample size.
So there is no number for how much duplicate entry costs you, and anybody quoting one is quoting something else. What the return on this work actually is, and what it would take to establish it, is a page of its own.
What to do first
Fix identifiers before anything visible. They are the join key, the error is permanent once goods ship, and every other improvement is easier once products can be told apart reliably.
Pull the party record out of your product rows. One list, keyed on the legal entity and on country and scheme. It is an afternoon and it stops the re-typing permanently.
Stop trying to reconcile the single-destination attributes. Decide which side of the line each one lives on and leave it there.
Treat composition as the exception worth doing properly. It is the one attribute with a genuine commerce destination and a genuine compliance destination, and it is the only place where the reuse argument is strong once the semantics are carried rather than the string. What the field is allowed to say and the four rules that already bind it are set out on fibre composition.
Do not buy a crosswalk. Buy, or build, the discipline of recording where each value came from. The mapping to each destination changes; the provenance does not. Running that over your own catalogue in a day is set out on the product data you already have.
What this page is not
It is not an argument for a product information platform. That is a decade-old category with funded incumbents and it solves a broader problem than this one. This page is about which apparel product attributes repay being fixed once, which is a narrower question.
It is not a data governance framework. The audience for those is engineering and the answer would be the wrong one.
It is not a claim that regulatory work fixes your feed. Most of what breaks a feed is commerce-only data that no instrument has any interest in.
What would change this page
A commerce or feed specification adding a home for recycled content, care instructions or a textile certification. That would move an attribute from the regulatory-only group into the universal one and change the sequencing above.
Any published measurement of duplicate attribute entry inside a single company. It would be the first, and it would let this page carry a number instead of an argument.
Sources
-
A cross-system attribute reuse analysis across seven product data systems, compiled 27 August 2026Field research
The load bearing source. Used for the four groups, for the two universal attributes, for the three false friends, for the commerce-only and regulatory-only sets, and for the finding that product category classification has no cross-system mapping.
-
A European Commission funded study on passport service provisionInstitutional research
Cited for one proposition: that potential providers were reluctant to supply quantitative figures for commercial reasons and in some cases did not have them.
-
A UK grocery product data quality report, October 2009Industry research
Cited only as the traceable origin of a widely restated figure, and for what it actually measured.
-
Search demand evidence on merchant product data language, compiled 27 August 2026Field research
Used for the merchant vocabulary in the opening and for the observation that this demand is served by feed tools, product information vendors and platform help centres rather than by anybody connecting it to product compliance. Demand in that research is qualitative throughout and no volume figure exists, so none appears here.
Sources as at 30 August 2026.
Keep going
The question this one usually raises next.
Also worth reading
- ImplementationThe product data you already have, and which of it a passport can useWhere each value actually lives, and the four questions that decide whether one can be published.
- Passport informationIdentifiers, and what allocating one commits you toThe join key, what a catalogue actually needs and what it costs.
Start with the data you already have.
No clean dataset required. ActivateDigital structures what you give it and shows what is still missing.
Help someone else make sense of product passports.