By: Johannes Fiegenbaum on 5/29/25, 7:15 AM · Last updated September 5, 2026
Scope 3 supply chain tracking is a data provenance problem before it is a software problem. The figure reported for purchased goods and services depends almost entirely on whether it came from a spend line multiplied by an industry emission factor, or from a supplier who calculated its own product footprint. No dashboard changes that. This guide sets out where the data comes from, what artificial intelligence genuinely solves, and the criteria worth applying before a tool is picked.
The GHG Protocol Corporate Value Chain (Scope 3) Standard sorts a value chain into 15 categories, and screening which of them are material comes before any data collection. Quality is a separate question. Every Scope 3 number sits in one of three tiers, and the tier decides what the number is worth. Spend-based data multiplies procurement spend by an industry emission factor: fast, complete, and blind to whether a supplier has decarbonised. Activity-based methods use physical quantities such as tonnes of steel or tonne-kilometres of freight combined with a generic factor, so they at least react to volume. Supplier-specific data comes from the supplier's own calculation, frequently routed through a CDP supply chain questionnaire, and is the only tier that rewards a supplier for actually cutting emissions.
The practical consequence is uncomfortable: a company can lower its reported Scope 3 by switching emission factor databases without changing anything physical.
| Data tier | What it uses | Reacts to supplier improvement | Effort | Audit evidence |
|---|---|---|---|---|
| Spend-based | Procurement spend times industry factor | No | Low | Ledger export plus factor source |
| Activity-based (average data) | Physical quantities times generic factor | Partly | Medium | Activity records from ERP or carrier |
| Supplier-specific | Supplier's own product or corporate footprint | Yes | High | Supplier calculation and assurance statement |
AI earns its place in two narrow jobs. The first is classification: mapping tens of thousands of procurement lines to categories and emission factors is tedious, largely rule-based, and genuinely faster with a trained model. The second is anomaly detection: flagging that a supplier's reported intensity moved by an order of magnitude between years.
What AI does not do is create primary data. A model can estimate a missing supplier footprint, but an assurance provider will treat that estimate as an estimate. Nor can it decide which factor set is appropriate for a category, a judgement call that moves results further than any algorithm does.
My position: continuous, sensor-level monitoring is the wrong first step for most mid-sized companies. Emission factor quality and supplier data coverage decide accuracy long before update frequency does, and a real-time feed sitting on spend-based inputs is an efficient way to be precisely wrong. Continuous monitoring also carries its own compute footprint. It is small next to the emissions being measured, but stating it is more honest than ignoring it.
Logistics is the one area where frequency pays off early. Transport activity data already exists in telematics and carrier systems as tonne-kilometres per lane, and carrier-specific factors are increasingly published, so freight emissions can be recalculated per shipment instead of waiting for an annual survey.
From building my own sustainability reporting software, the recurring gap is between methods that are correct on paper and methods that survive contact with a real supplier list. Five criteria separate the two, and none of them appear on a feature comparison page.
| Criterion | What to ask | Failure mode in practice |
|---|---|---|
| Factor library provenance | Which database, which version, updated when? | Year-on-year comparisons break silently when the library updates mid-cycle |
| Supplier data intake | Can a supplier submit a figure without buying a licence? | Coverage stalls at the suppliers willing to create an account |
| ERP integration | Does spend and material data flow automatically or by CSV upload? | The number is only ever as fresh as the last manual export |
| Audit trail | Can every figure be traced to a source document and a method? | Assurance turns into an archaeology project |
| Recalculation | Can the base year be restated when method or perimeter changes? | Targets stop meaning anything after the first acquisition |
Recalculation is the criterion most often skipped and the one that hurts most: a setup that cannot restate a base year forces a later choice between an honest number and a comparable one. The same goes for integration: the point of connecting sustainability data through APIs is not the dashboard, it is that nobody has to remember to export anything.
Sequence matters more than tooling. Rank suppliers by spend first, because that list exists on day one, then re-rank by estimated emissions once a spend-based screening is in place. In most portfolios a small share of suppliers carries the bulk of upstream emissions, and only those are worth a bespoke data request in the first year.
The minimum data set to ask a supplier for is short:
Anything longer depresses response rates without improving the number. What an external engagement realistically delivers in the first 90 days is a complete spend-based baseline across the relevant categories, a ranked hotspot list, a documented method, and a request pack for the top tier of suppliers. Primary data arrives in the following quarter at the earliest, because suppliers need a reporting cycle of their own to answer at all. By month twelve the realistic outcome is supplier-specific data for the largest suppliers, estimates for the rest, and a base year that can carry science-based Scope 3 targets. An SBTi submission stands or falls on that base year, not on the reporting layer built around it.
ESRS E1-6 asks for gross Scope 3 greenhouse gas (GHG) emissions by significant category, together with the calculation method, the emission factors used and the share of primary data behind the figure. The disclosure is about provenance as much as magnitude: an estimate labelled as an estimate is compliant, an undocumented number is not. That is precisely why the tier table above matters more than update frequency, and it is the same logic that drives climate risk disclosure and the VSME standard smaller suppliers report against.
The report corpus I work with shows how uneven the practice still is. Across 1,401 public 2024 and 2025 reports from European listed companies, 69 percent contain complete Scope 1, 2 and 3 data, 13 percent publish no extractable Scope data at all, and 8 percent of those that do report a Scope 3 figure show it as smaller than Scope 1 or 2, which is a methodological red flag rather than an achievement. The bottleneck in those reports is not missing software. It is a value chain nobody has asked yet.
How this plays out in practice: Corporate PPAs in Europe: Offtake Tenders, Structures and Reporting.
Which leads to the one thing worth deciding before any budget is spent: Scope 3 data assembled to satisfy a reporting deadline reliably produces a conformant report and nothing a buyer can act on. The test that separates the two is blunt. If a category figure cannot change a sourcing decision, a supplier conversation or a target, it is a reporting number, not a steering number. Both are legitimate outputs, but only one of them justifies a tool.
Because the data belongs to other companies. Scope 1 and 2 come from your own meters and invoices, while Scope 3 depends on suppliers and carriers with their own systems and cycles. The difficulty is organisational, not computational.
Under the European Sustainability Reporting Standards they are mandatory where material, disclosed by significant category under E1-6. Outside that scope, Scope 3 reporting is driven by customer requirements and CDP supply chain requests, not by law.
Data is good enough when it ranks hotspots correctly and is documented well enough to be restated later. Chasing supplier-specific figures for categories that cannot change the ranking is how a programme stalls in year one.
Your upstream emissions are somebody else's Scope 1, so the same tonne is counted several times along a chain. That is inherent to value chain accounting, not an error. It only misleads when several parties claim the same reduction separately.
ESG and sustainability consultant based in Hamburg, specialised in VSME reporting and climate risk analysis. Has supported 300+ projects for companies and financial institutions, from mid-sized manufacturers to major banks and insurers.
More aboutCarbon accounting fails in young companies for reasons that have little to do with climate science. The numbers exist somewhere: in invoices, in travel bookings, in a supplier's ...
Read more →