Skip to content

Copiara is coming to the Shopify App Store. It is not listed yet.

Copiara
Back to the journal
Product data8 min read

Product Attribute Completeness for B2B Search

Product attribute completeness for B2B search means every SKU carries the attributes buyers filter on in its category. How to find and close the gaps.

By Amir Hessabi

Product attribute completeness for B2B search is the share of SKUs in a category that carry a value for every attribute buyers filter and compare on in that category. It is measured per category, not across the whole catalog, because a bearing and a circuit breaker are found by completely different attributes. A SKU missing one filterable attribute does not rank a little lower in B2B search. It disappears from the filtered result.

Most distributors hear "fix your product data" and picture a multi-year cleanup of every field on every SKU. That project never gets approved, and it should not. The job is much smaller and much more specific than that, and this post is about how to scope it.

It is written for the operations or ecommerce lead who owns the catalog on a distributor's Shopify store. You do not need to be a developer to do any of it.

Why does one blank attribute hide a product a buyer wanted?#

Take a 6205-2RS deep-groove ball bearing, one of the most common bearings in any MRO catalog.

A maintenance buyer who needs one does not usually type a brand. They filter. Bore 25 mm, outside diameter 52 mm, width 15 mm, sealed both sides, C3 clearance because the motor runs hot. Five attributes, and they are the whole search.

Now suppose your SKU has bore, OD, width, and seal type filled in, and the clearance field is blank. The buyer ticks C3. Your bearing drops out of the result. The buyer sees another brand's 6205, or they see nothing and pick up the phone, or they order from whoever did fill in the field.

The description does not rescue it. Your product copy can say "C3 clearance" in plain English, and it will not matter, because filters and comparison tables read structured fields, not prose. A blank field is not "partially complete" to a filter. It is a no.

The same mechanic applies anywhere a machine compares products. Distribution Strategy Group put it plainly on September 16: AI-powered search and purchasing tools use specifications, dimensions, and other attributes to identify and compare products, and incomplete or inconsistent data limits their ability to return useful results. Whether the thing doing the filtering is a checkbox or a model, the missing field is invisible.

Which attributes actually matter in each category?#

Here is the part that makes the job small. Most categories have a short list, often somewhere around four to eight attributes, that do the actual finding. The manufacturer's spec sheet may list forty. Buyers filter on a handful.

Split every category's attributes into two tiers:

  • Findability attributes. The ones buyers filter on and compare across. For the bearing: bore, OD, width, seal or shield type, clearance. For a molded case circuit breaker, it is a different list entirely: amperage, poles, voltage rating, interrupting rating, frame. These are the attributes your completeness target applies to.
  • Reference attributes. Everything else on the datasheet: weight, dynamic load rating, limiting speed, country of origin. Useful on the product page, rarely the reason a buyer finds the SKU.

Do not pick the findability list from the spec sheet. Pick it from how buyers actually ask. Read a week of phone and email requests for the category and write down which attributes people name. Then look at your site search log for the same category and see which values people type. If buyers keep saying "sealed, C3," those are findability attributes, whatever the datasheet puts first.

If search logs are new territory, our post on failed site search analysis for distributors covers how to read them. This post assumes you already know what buyers are looking for and asks a narrower question: is the field there when they look?

How do you measure completeness without boiling the ocean?#

A per-category scorecard is enough. One row per category, five columns:

CategorySKUsFindability attributesSKUs with all of themMost-missing attribute
Deep-groove ball bearingscountbore, OD, width, seal, clearancecount and percentthe field to fix first
Molded case circuit breakerscountamps, poles, voltage, interrupting rating, framecount and percentthe field to fix first
Hex cap screwscountdiameter, thread pitch, length, grade, finishcount and percentthe field to fix first

The last column is the useful one. A category is often not missing everything. It may be missing one attribute across a lot of SKUs, for example because the feed that populated the category never carried it. Find that attribute and you have found the cheapest fix in the catalog.

Two rules keep this from turning into the multi-year project:

  • Weight by demand. Start with the categories that sell and the categories people search for. A category you ship twice a year can wait. Do not work alphabetically.
  • Do not chase a catalog-wide number. A single completeness percentage across every SKU averages away exactly the gaps that matter. A catalog can look healthy overall while its best-selling category is invisible to half its filters.

We are not going to give you an industry-average completeness figure to aim at, because we have not seen a trustworthy one. The target that matters is yours: every findability attribute, filled, on the categories buyers shop.

Where do the missing values come from, and how do you keep track?#

Filling a field is easy. Filling it with a value you can stand behind, and knowing later where it came from, is the actual discipline.

Rank your sources by how much you trust them:

  1. Manufacturer feeds and industry data exchanges. Structured, maintained by the people who make the part.
  2. Manufacturer datasheets. Authoritative, but someone has to read them and type the value in.
  3. Your own staff's knowledge. The counter veteran who knows every 6205 you stock runs C3. Often right, rarely written down.
  4. Inferred or generated values. Parsed from a description, or suggested by a tool. Useful as a placeholder, never the last word.

The supply side of the first tier just moved. On September 16, 2026, Distribution Strategy Group reported that Salsify and the Industry Data Exchange Association (IDEA) launched an integration letting electrical manufacturers send product data straight from Salsify's PIM into IDEA Connector, with one prebuilt channel for core product data and a second for category-specific attributes. It is in an early access program for Salsify customers, and neither company disclosed how many manufacturers are participating or how much time it saves.

For a distributor, the takeaway is not about either vendor. It is that better category-specific attribute data is starting to flow from manufacturers, and your catalog only benefits if it can accept a better value when one arrives. That requires two things:

  • Provenance on every value. Record where each attribute came from and when it changed. A clearance value typed in from memory and a clearance value from the manufacturer's feed should not look identical in your system.
  • A rule for who wins. When a manufacturer value arrives for a field your team filled with a placeholder, the better source should take over without a fight. When two good sources disagree, a person should look at it. When someone on your team has deliberately set a value, an automated import should not quietly overwrite it.

Without those, every new feed is either a risk (it overwrites good values) or a waste (nobody trusts it enough to load it).

How does Copiara approach attribute completeness on a Shopify catalog?#

Copiara is a Shopify app for wholesale distributors, currently in build and not yet listed on the Shopify App Store, so nobody can install it today. Catalog enrichment is one of its modules, and it works the way this post argues a catalog should.

It layers attributes, provenance, and field-level change history over your Shopify products without overwriting them, plus readiness scoring. Shopify stays the system of record for its own product fields. Imports enrich rather than blank: an incoming empty value never wipes a populated field. Every product change writes a field-level revision, so you can see what changed, from which source, and when. Sources carry a precedence rank, so a real feed displaces a lower-ranked generated value, two real sources that disagree go to a human review queue instead of silently swapping, and a manual lock outranks everything.

Copiara is in early access. If attribute gaps are costing you searches, get in touch here.

Frequently asked questions#

Is 100 percent completeness the goal? No. The goal is complete findability attributes on the categories buyers actually shop. Reference attributes can lag behind without costing you a search.

Can generated or inferred values fill the gaps? They can hold a place, as long as they are labeled with their source and yield to a manufacturer value when one arrives. An unlabeled guess that looks exactly like a verified value is worse than a blank.

Does Copiara overwrite my Shopify product data? No. Enrichment is layered over your Shopify products. Shopify remains the system of record for its own fields.

Do I need a PIM to do this? What you need is structured attributes per category, with a record of where each value came from. How you get there depends on the size of your catalog and your team. Start with the scorecard; it will tell you how big the problem really is.

Before the listing goes live

See Copiara on your own Shopify store.

If cross-referencing, negotiated quotes, or buyer approvals sound like your buyers' problem, get on the early access list.