Skip to content
Build logDigital transformation

How we build product pages from manufacturer documents: a sourced comparison table on 100+ pages

A product page is a release, not a text. We build it from the manufacturer’s documents through a fixed structure, a quality gate and a sign-off, so a launch need not wait for a writer.

Contents
  1. The commercial question
  2. Why a pipeline
  3. How we worked
  4. A value without a source does not ship
  5. The loop
  6. One example
  7. What a slow launch costs
  8. Scaling to a second channel
  9. Guardrails
  10. Steering
  11. What it unlocks
  12. What we would tell another business

An online shop sells technical products from several manufacturers, and the facts its buyers need sat in manuals, spec sheets and brochures. We built a pipeline that turns those documents into product pages with one structure and a comparison table.

The pipeline now builds more than 100 of the shop’s pages with a comparison table. We count each shop listing as one page, including bundles, variants and B-stock. In the launch described below, new models went from a fresh ERP record to checked live pages in one working session. The same tables also supply the import data for eBay, a second channel.

What the pipeline changes on each product page.

  • Where the facts come from
    A value without a source comes out
    Each table value traces to a source document
    Before
    Written or pasted page by page
    After
    Manufacturer documents from official sites, permitted products only
  • A new model’s page
    New records are checked for inherited titles first
    The launch does not wait for a writer
    Before
    Waits for a writer
    After
    Built from the documents, checked, signed off and imported
  • A second channel
    First import set delivered as files
    The table is written once
    Before
    No shared product data
    After
    eBay item specifics and titles read from the same table
The pipeline now builds more than 100 of the shop’s pages, counting bundles, variants and B-stock.

The commercial question

A product page earns money through drivers that you can move. You get more visitors when structured data travels to eBay and other channels. Conversion rises when a buyer can compare models and pick the right one. Returns and support calls fall when the specifications and box contents are accurate. And you gain selling days when a new model’s page is live as the model arrives.

What mattered was how each page could carry the manufacturer’s facts, and how fast a new model could start selling. Since 13 December 2024, the EU’s General Product Safety Regulation also sets what an online offer has to show. That includes the manufacturer with a postal and electronic address, and a responsible person in the EU where the manufacturer is based outside it. It also includes the product’s identification with a picture, and warnings or safety information. Our quality gate checks that each page carries the manufacturer’s contact block and an EU responsible person. It does not check warnings or safety texts, so passing it is not a compliance review.

Value driver tree: how a product page earns Contribution from a product equals visitors times conversion times price minus product cost, less returns and support cost, over the days it is on sale. Page levers: structured data and more channels raise visitors, a sourced comparison table raises conversion because buyers find the right model, accurate specifications and box contents reduce wrong purchases and returns, and launch speed adds selling days. How a product page earns money The page moves four of the five drivers. Product contribution over its time on sale Visitors more channels, structured data same table feeds eBay Conversion buyer finds the right model sourced comparison table Price, product cost set by the business not moved by the page Returns, support fewer wrong purchases accurate specs, box contents Selling days live when the model arrives one-session launch Page levers Teal = moved by the product page and its data. Grey = set elsewhere. Contribution = visitors × conversion × (price − product cost) − returns and support, over selling days. Value driver tree: how a product page earns Contribution from a product equals visitors times conversion times price minus product cost, less returns and support cost, over the days it is on sale. Page levers: structured data and more channels raise visitors, a sourced comparison table raises conversion because buyers find the right model, accurate specifications and box contents reduce wrong purchases and returns, and launch speed adds selling days. How a product page earns money The page moves four of the five drivers. Product contribution over its time on sale Visitors more channels, structured data same table feeds eBay Conversion buyer finds the right model sourced comparison table Price, product cost set by the business not moved by the page Returns, support fewer wrong purchases accurate specs, box contents Selling days live when the model arrives one-session launch Page levers Teal = moved by the product page and its data. Grey = set elsewhere. Contribution = visitors × conversion × (price − product cost) − returns and support, over selling days.
How a product page earns money. The page moves four of the five drivers.

Why a pipeline

We weighed the ways to fill the pages by what each page costs and how fast it goes live.

Pages written one at a time can read well, yet each one costs the same, launches wait for a writer, and the facts depend on who checked them. Pasting the manufacturer’s text is fast and cheap, and it leaves pages thin for a buyer, often in the wrong language, with nothing that compares one model with the next.

A governed pipeline from manufacturer documents needs a build and a set of rules first. After that, one more page costs little, a new model’s page does not wait for a writer, and the same table feeds eBay.

The pipeline scales with the range, and that decided it. We automated the routine and put a person at every gate that matters.

How we worked

We worked inside the client’s ERP, JTL-Wawi, and its shop. The ERP stays the master for product data, and we read it through read-only snapshots. Changes reach the shop only through the ERP’s own import tool, and only after sign-off, and the ERP’s connector then publishes them.

We asked the manufacturers for permission to use their content, and we build pages from it only for the products they permitted. Documents came from the manufacturers’ official websites and nowhere else. We collected manuals, spec sheets, brochures, images and videos. Where a site exists in several languages, we matched products across all language versions.

Existing pages are protected by design: pages already in the new format are locked at several separate points, and existing attributes are never overwritten by default. A deliberate correction needs its own file and its own approval.

The ERP article is the hub of the data model. Products and their source documents sit on one side. Attributes, the description with its table and the channel listings hang off the article.

A value without a source does not ship

A comparison table invites trust, because each cell looks like a fact. So each cell has to trace back to a manufacturer document.

When we rebuilt the tables for several product families against the manufacturers’ manuals, the check changed real values. Several values in one table did not match the manual, and a merged setting had to be split. One claim had no source in this product’s documents, so the row came out.

Source tracing: a table value ships only if the manufacturer's document states it Based on real corrections. Left, five rows of a comparison table with the draft value and the checked value. Three values were corrected to what the manufacturer's manual states, a merged setting was split into two rows, and a speed claim with no source in this product's documents was removed. Right, the manufacturer's manual with its technical specifications block. Teal lines connect confirmed values to the document. An amber line leads to no source found, row removed. A value ships only if the manufacturer's document says it Based on real corrections. Table row Draft Shipped Runtime, unit A 12 h 8 to 14 h Range, unit B 20 m 15 m Runtime, unit C 9 h about 7 h Combined setting merged in one two separate rows Speed claim 3x faster row removed Manufacturer's manual Technical specifications block Official manufacturer site only No source in this product's documents Where the documents give no value, the row comes out. We do not fill it with a plausible guess. Rule: every table cell must be found in the curated manufacturer documents. Trace is kept at document level, not per cell. Source tracing: a table value ships only if the manufacturer's document states it Based on real corrections. Left, five rows of a comparison table with the draft value and the checked value. Three values were corrected to what the manufacturer's manual states, a merged setting was split into two rows, and a speed claim with no source in this product's documents was removed. Right, the manufacturer's manual with its technical specifications block. Teal lines connect confirmed values to the document. An amber line leads to no source found, row removed. A value ships only if the manufacturer's document says it Based on real corrections. Table row Draft Shipped Runtime, unit A 12 h 8 to 14 h Range, unit B 20 m 15 m Runtime, unit C 9 h about 7 h Combined setting merged in one two separate rows Speed claim 3x faster row removed No source in this product's documents Manufacturer's manual Technical specifications block Official manufacturer site only Where the documents give no value, the row comes out. We do not fill it with a plausible guess. Rule: every table cell must be found in the curated manufacturer documents. Trace is kept at document level, not per cell.
A value ships only if the manufacturer’s document says it. Based on real corrections, with neutral row names.

The kinds of correction in this chart are real. The rule since then is simple: a table value has to be found in the curated manufacturer documents. Where the documents give no value, we remove the row rather than fill it with a plausible guess. We keep the trace at the level of the document, not the single cell. A reviewer can find the document behind a value, but not the page within it.

The loop

Each product runs the same eight steps, from PDF to live page.

  1. Harvest. Collect documents and media from the manufacturer’s official site.
  2. Extract. Read the PDF text, with OCR for scanned files, into one record per product.
  3. Compose. An AI system drafts the page text and its table from the product’s record, in the fixed structure. B-stock pages are derived from the main page without it.
  4. Check. A quality gate tests the structure and a duplicate guard cleans repeated blocks. The gate does not test facts. Failing pages are held back.
  5. Approve. A person reviews each page on an approval sheet and signs the batch. Whether the facts are right rests on this review.
  6. Import. The ERP’s import tool loads master data, attributes and function attributes in three runs.
  7. Verify. We read the live product feed and confirm each page carries its table.
  8. Reuse. The same tables feed the eBay import.
Release process with decision points Row one: harvest documents from the official manufacturer site, extract text with OCR fallback, compose the page in the canonical structure with the table from the documents, then decision one: does the page pass the structure gate? If no, it is held back and returns to compose. Row two: if yes, it goes on the approval sheet, decision two: signed? If no, it returns to compose. If yes, decision three: does the ERP record carry inherited values? If yes, a targeted override file with its own approval. Then compile import files. Row three: three import runs for master data, attributes and function attributes, the ERP connector syncs to the shop, decision four: is the table live on the page? If no, fix and re-run from compose. If yes, done, and the eBay export reads the same tables for item specifics and titles. Release process with decision points Teal is the path a page takes when every check passes. Harvest official site only Extract text, OCR fallback Compose canonical, table Passes the structure gate? Held back not imported no fix and compose again yes Approval sheet tick per listing Signed? no: revise ERP record carries inherited values? Override own file, approval yes Compile imports two attribute files no Import runs x3 master data, attributes Connector sync ERP to shop Table live on the page? Done live, verified yes no: fix and re-run eBay export same tables Source type: pipeline documentation and launch log. yes Release process with decision points Row one: harvest documents from the official manufacturer site, extract text with OCR fallback, compose the page in the canonical structure with the table from the documents, then decision one: does the page pass the structure gate? If no, it is held back and returns to compose. Row two: if yes, it goes on the approval sheet, decision two: signed? If no, it returns to compose. If yes, decision three: does the ERP record carry inherited values? If yes, a targeted override file with its own approval. Then compile import files. Row three: three import runs for master data, attributes and function attributes, the ERP connector syncs to the shop, decision four: is the table live on the page? If no, fix and re-run from compose. If yes, done, and the eBay export reads the same tables for item specifics and titles. Release process with decision points Teal is the path a page takes when every check passes. Harvest official site only Extract text, OCR fallback Compose canonical, table Passes the structure gate? Held back not imported no fix and compose again yes Approval sheet tick per listing Signed? no: revise yes ERP record carries inherited values? no Compile imports two attribute files Override own file, approval yes Import runs x3 master data, attributes Connector sync ERP to shop Table live on the page? no: fix and re-run yes Done live, verified eBay export same tables Source type: pipeline documentation and launch log.
Release process with decision points. Teal is the path a page takes when every check passes.

One example

New models are often created in the ERP by copying the previous generation’s record. That is the fastest way to create a record. The copy carries old titles, and a never-overwrite rule then protects the wrong values.

working-notes · thread

  1. Team

    The new records are in the ERP. Can we compile the import?

  2. Capcelerate systemSystem

    Not yet. The meta titles still carry the previous generation’s model numbers. The never-overwrite rule would keep them.

  3. Team

    Override those two fields only. Everything else stays protected.

  4. Capcelerate systemSystem

    A separate attribute file, two fields per page. It is on the approval sheet for sign-off.

Recreated from our working notes. Details anonymised.

The table for the new line came from the manufacturer’s series brochure and covered four models. After sign-off the imports ran, and we read the live product feed. All four pages were live with their full comparison tables within one working session.

The fix became a step, and before each launch we now check new records for inherited titles and an inherited web address.

What a slow launch costs

A new model sells only once its page is live. Each day without a page is a day of demand that goes elsewhere.

Sales lost to a late product page Line chart of cumulative units sold over the first 60 days after a new model arrives, at 3 units a day once the page is live. Page live on day 1: 177 units by day 60. Page live on day 22: 114 units. The shaded gap of 63 units is the launch window lost to a slow page. What a slow launch costs Cumulative units in the first 60 days of a new model. The page decides when selling starts. 0 50 100 150 200 day 0 day 10 day 20 day 30 day 40 day 50 day 60 page live day 1: 177 day 22: 114 63 units never sold Assumptions: 3 units a day once the page is live, a page three weeks late (day 22), no catch-up later. Sales lost to a late product page Line chart of cumulative units sold over the first 60 days after a new model arrives, at 3 units a day once the page is live. Page live on day 1: 177 units by day 60. Page live on day 22: 114 units. The shaded gap of 63 units is the launch window lost to a slow page. What a slow launch costs Cumulative units in the first 60 days of a new model. The page decides when selling starts. page live day 1: 177 day 22: 114 63 units never sold 0 50 100 150 200 day 0 day 20 day 40 day 60 Assumptions: 3 units a day once the page is live, a page three weeks late (day 22), no catch-up later.
What a slow launch costs. Cumulative units in the first 60 days of a new model. Assumes 3 units a day once the page is live.

The example assumes a page that goes live on day 22 instead of day 1. Under it, the late page loses about a third of the first two months’ units. The example also assumes that no buyer waits for the late model. Buyers who want one specific model often do wait, so a real loss can be smaller. The pipeline is built to keep that gap short, and in the launch above the pages went live within one working session.

Scaling to a second channel

We write the table once, and each channel reads from it.

One sourced comparison table, three uses Centre, one comparison table built from manufacturer documents. Arrows lead to three uses: the shop page shows the table, eBay item specifics take one row per feature with the value from the highlighted product column, and eBay titles of at most 80 characters are built only from features in the table with no duplicate titles. The shop page and the eBay import files read from the same table. One sourced table, three uses The table is written once. Each channel reads from it. Comparison table Values from the manufacturer's documents this product Shop page The table sits on the product page, same section on every page eBay item specifics One row per feature, value taken from the highlighted column eBay titles Generated only from table features, at most 80 characters, no duplicates The shop page and every eBay import file read from the same tables. Import files prepared through the ERP. Titles checked for length and for duplicates across listings. One sourced comparison table, three uses Centre, one comparison table built from manufacturer documents. Arrows lead to three uses: the shop page shows the table, eBay item specifics take one row per feature with the value from the highlighted product column, and eBay titles of at most 80 characters are built only from features in the table with no duplicate titles. The shop page and the eBay import files read from the same table. One sourced table, three uses The table is written once. Each channel reads from it. Comparison table Values from the manufacturer's documents this product Shop page The table sits on the product page, same section on every page eBay item specifics One row per feature, value taken from the highlighted column eBay titles Generated only from table features, at most 80 characters, no duplicates The shop page and every eBay import file read from the same tables. Import files prepared through the ERP. Titles checked for length and for duplicates across listings.
One sourced table feeds the shop page, eBay item specifics and eBay titles.

For eBay, the same descriptions are reused without the video and manufacturer sections. Item specifics come straight from the table, with one row per feature and the value for this product. Where the client has not approved a title, we generate one only from table features. It stays within 80 characters and is never duplicated. The first eBay import set went to the client as files.

Scaling surfaced finds of its own:

  • A channel filter, corrected. An early eBay export read another channel’s setting and left out products that belonged in the set. With the right setting, the next set included them.
  • Images reused, not uploaded twice. Images and videos already on the shop’s earlier pages were carried over into the new pages rather than uploaded again.
  • Manuals checked by fingerprint. We check manual links by file fingerprint, and a manual in the wrong language gets flagged for correction.
  • Re-runs that should change nothing. In early runs, some bundle pages showed a block several times. Each step now replaces its block, and a guard removes repeated blocks before each import file is written.

Guardrails

GuardrailWhat it does
Official sources onlyNew documents, images and videos come from the manufacturers’ official sites, and only for products the manufacturer permitted. Never from distributors or third-party sites. Higher resolution has to come from the manufacturer.
Fixed structureA normaliser builds every page in the same section order.
Quality gateTests each page for the required sections, the table and the manufacturer block with its contact details. It does not test facts. Failing pages are held back.
Duplicate guardRemoves repeated blocks before any import file is written.
Human sign-offEach run produces an approval sheet. Without a tick and a signature, nothing enters the import file.
Locked pagesFinished pages are protected at several points. Attributes are only filled where empty.
Record check before launchNew records are checked for inherited titles and web addresses before the import is compiled.
Split importsAttributes and function attributes go in separate files. In our runs, a file that mixed both imported only half its rows.
Live checkThe live product feed is read after every import.

One incident shows the review working on our own work. For a launch, the first draft was written outside the pipeline with its own prose and its own table. We rejected it and rebuilt the pages through the pipeline. Since then the rule is written down: no new page without the pipeline.

Where a manufacturer only offers low-resolution images, we place them on a larger white canvas. We do not take copies from other sellers.

Steering

The standard was our own, and our project lead at Capcelerate set it in one sentence.

“Everything must be built exactly the same way.”

Capcelerate project lead, translated from German

We ruled on the trade-offs behind it. Where a value had no source, we removed the row rather than print “not available”, which reads cleaner for a buyer. Where an existing hand-edited live page was richer than the pipeline’s version, we kept the live page.

What it unlocks

Once the data is built, a later channel or launch reuses it and does not go back to the documents.

Accessories need their own template, while other product families follow the same pipeline. Pages with a table and pages without one can be compared on conversion, returns and support contacts. The next euro can then go where pages earn most.

What we would tell another business

  1. Treat a product page as a release. It needs a structure, a source, a gate, a sign-off and a live check.
  2. Source each fact from the manufacturer. Remove what you cannot source rather than guess.
  3. Write the table once and publish it to each channel. Channels then read from one record.
  4. Automate the routine and gate the exceptions. Your people sign off rather than retype.
  5. Protect what is already good. Never overwrite a finished page by default.

If your product pages say less than your manufacturers’ documents, we can start with one product family.

Credits

Built by the Capcelerate team. Thanks to the client’s team, who created the new records and ran the launch imports.

Tell us what you want to grow. The intro call is free.

Request a free intro callFree, up to 30 minutes. No obligation.