An online shop sells technical products from several manufacturers, and the facts its buyers need sat in manuals, spec sheets and brochures. We built a pipeline that turns those documents into product pages with one structure and a comparison table.
The pipeline now builds more than 100 of the shop’s pages with a comparison table. We count each shop listing as one page, including bundles, variants and B-stock. In the launch described below, new models went from a fresh ERP record to checked live pages in one working session. The same tables also supply the import data for eBay, a second channel.
What the pipeline changes on each product page.
-
Where the facts come fromA value without a source comes outEach table value traces to a source document
- Before
- Written or pasted page by page
- After
- Manufacturer documents from official sites, permitted products only
-
A new model’s pageNew records are checked for inherited titles firstThe launch does not wait for a writer
- Before
- Waits for a writer
- After
- Built from the documents, checked, signed off and imported
-
A second channelFirst import set delivered as filesThe table is written once
- Before
- No shared product data
- After
- eBay item specifics and titles read from the same table
The commercial question
A product page earns money through drivers that you can move. You get more visitors when structured data travels to eBay and other channels. Conversion rises when a buyer can compare models and pick the right one. Returns and support calls fall when the specifications and box contents are accurate. And you gain selling days when a new model’s page is live as the model arrives.
What mattered was how each page could carry the manufacturer’s facts, and how fast a new model could start selling. Since 13 December 2024, the EU’s General Product Safety Regulation also sets what an online offer has to show. That includes the manufacturer with a postal and electronic address, and a responsible person in the EU where the manufacturer is based outside it. It also includes the product’s identification with a picture, and warnings or safety information. Our quality gate checks that each page carries the manufacturer’s contact block and an EU responsible person. It does not check warnings or safety texts, so passing it is not a compliance review.
Why a pipeline
We weighed the ways to fill the pages by what each page costs and how fast it goes live.
Pages written one at a time can read well, yet each one costs the same, launches wait for a writer, and the facts depend on who checked them. Pasting the manufacturer’s text is fast and cheap, and it leaves pages thin for a buyer, often in the wrong language, with nothing that compares one model with the next.
A governed pipeline from manufacturer documents needs a build and a set of rules first. After that, one more page costs little, a new model’s page does not wait for a writer, and the same table feeds eBay.
The pipeline scales with the range, and that decided it. We automated the routine and put a person at every gate that matters.
How we worked
We worked inside the client’s ERP, JTL-Wawi, and its shop. The ERP stays the master for product data, and we read it through read-only snapshots. Changes reach the shop only through the ERP’s own import tool, and only after sign-off, and the ERP’s connector then publishes them.
We asked the manufacturers for permission to use their content, and we build pages from it only for the products they permitted. Documents came from the manufacturers’ official websites and nowhere else. We collected manuals, spec sheets, brochures, images and videos. Where a site exists in several languages, we matched products across all language versions.
Existing pages are protected by design: pages already in the new format are locked at several separate points, and existing attributes are never overwritten by default. A deliberate correction needs its own file and its own approval.
The ERP article is the hub of the data model. Products and their source documents sit on one side. Attributes, the description with its table and the channel listings hang off the article.
A value without a source does not ship
A comparison table invites trust, because each cell looks like a fact. So each cell has to trace back to a manufacturer document.
When we rebuilt the tables for several product families against the manufacturers’ manuals, the check changed real values. Several values in one table did not match the manual, and a merged setting had to be split. One claim had no source in this product’s documents, so the row came out.
The kinds of correction in this chart are real. The rule since then is simple: a table value has to be found in the curated manufacturer documents. Where the documents give no value, we remove the row rather than fill it with a plausible guess. We keep the trace at the level of the document, not the single cell. A reviewer can find the document behind a value, but not the page within it.
The loop
Each product runs the same eight steps, from PDF to live page.
- Harvest. Collect documents and media from the manufacturer’s official site.
- Extract. Read the PDF text, with OCR for scanned files, into one record per product.
- Compose. An AI system drafts the page text and its table from the product’s record, in the fixed structure. B-stock pages are derived from the main page without it.
- Check. A quality gate tests the structure and a duplicate guard cleans repeated blocks. The gate does not test facts. Failing pages are held back.
- Approve. A person reviews each page on an approval sheet and signs the batch. Whether the facts are right rests on this review.
- Import. The ERP’s import tool loads master data, attributes and function attributes in three runs.
- Verify. We read the live product feed and confirm each page carries its table.
- Reuse. The same tables feed the eBay import.
One example
New models are often created in the ERP by copying the previous generation’s record. That is the fastest way to create a record. The copy carries old titles, and a never-overwrite rule then protects the wrong values.
working-notes · thread
-
Team
The new records are in the ERP. Can we compile the import?
-
Capcelerate systemSystem
Not yet. The meta titles still carry the previous generation’s model numbers. The never-overwrite rule would keep them.
-
Team
Override those two fields only. Everything else stays protected.
-
Capcelerate systemSystem
A separate attribute file, two fields per page. It is on the approval sheet for sign-off.
The table for the new line came from the manufacturer’s series brochure and covered four models. After sign-off the imports ran, and we read the live product feed. All four pages were live with their full comparison tables within one working session.
The fix became a step, and before each launch we now check new records for inherited titles and an inherited web address.
What a slow launch costs
A new model sells only once its page is live. Each day without a page is a day of demand that goes elsewhere.
The example assumes a page that goes live on day 22 instead of day 1. Under it, the late page loses about a third of the first two months’ units. The example also assumes that no buyer waits for the late model. Buyers who want one specific model often do wait, so a real loss can be smaller. The pipeline is built to keep that gap short, and in the launch above the pages went live within one working session.
Scaling to a second channel
We write the table once, and each channel reads from it.
For eBay, the same descriptions are reused without the video and manufacturer sections. Item specifics come straight from the table, with one row per feature and the value for this product. Where the client has not approved a title, we generate one only from table features. It stays within 80 characters and is never duplicated. The first eBay import set went to the client as files.
Scaling surfaced finds of its own:
- A channel filter, corrected. An early eBay export read another channel’s setting and left out products that belonged in the set. With the right setting, the next set included them.
- Images reused, not uploaded twice. Images and videos already on the shop’s earlier pages were carried over into the new pages rather than uploaded again.
- Manuals checked by fingerprint. We check manual links by file fingerprint, and a manual in the wrong language gets flagged for correction.
- Re-runs that should change nothing. In early runs, some bundle pages showed a block several times. Each step now replaces its block, and a guard removes repeated blocks before each import file is written.
Guardrails
| Guardrail | What it does |
|---|---|
| Official sources only | New documents, images and videos come from the manufacturers’ official sites, and only for products the manufacturer permitted. Never from distributors or third-party sites. Higher resolution has to come from the manufacturer. |
| Fixed structure | A normaliser builds every page in the same section order. |
| Quality gate | Tests each page for the required sections, the table and the manufacturer block with its contact details. It does not test facts. Failing pages are held back. |
| Duplicate guard | Removes repeated blocks before any import file is written. |
| Human sign-off | Each run produces an approval sheet. Without a tick and a signature, nothing enters the import file. |
| Locked pages | Finished pages are protected at several points. Attributes are only filled where empty. |
| Record check before launch | New records are checked for inherited titles and web addresses before the import is compiled. |
| Split imports | Attributes and function attributes go in separate files. In our runs, a file that mixed both imported only half its rows. |
| Live check | The live product feed is read after every import. |
One incident shows the review working on our own work. For a launch, the first draft was written outside the pipeline with its own prose and its own table. We rejected it and rebuilt the pages through the pipeline. Since then the rule is written down: no new page without the pipeline.
Where a manufacturer only offers low-resolution images, we place them on a larger white canvas. We do not take copies from other sellers.
Steering
The standard was our own, and our project lead at Capcelerate set it in one sentence.
“Everything must be built exactly the same way.”
We ruled on the trade-offs behind it. Where a value had no source, we removed the row rather than print “not available”, which reads cleaner for a buyer. Where an existing hand-edited live page was richer than the pipeline’s version, we kept the live page.
What it unlocks
Once the data is built, a later channel or launch reuses it and does not go back to the documents.
Accessories need their own template, while other product families follow the same pipeline. Pages with a table and pages without one can be compared on conversion, returns and support contacts. The next euro can then go where pages earn most.
What we would tell another business
- Treat a product page as a release. It needs a structure, a source, a gate, a sign-off and a live check.
- Source each fact from the manufacturer. Remove what you cannot source rather than guess.
- Write the table once and publish it to each channel. Channels then read from one record.
- Automate the routine and gate the exceptions. Your people sign off rather than retype.
- Protect what is already good. Never overwrite a finished page by default.
If your product pages say less than your manufacturers’ documents, we can start with one product family.
Credits
Built by the Capcelerate team. Thanks to the client’s team, who created the new records and ran the launch imports.