20–25 Sept 2026
Aalborg University & Online
Europe/Copenhagen timezone

PCFBench: an open, task-level diagnostic benchmark for AI-assisted product carbon footprinting

Not scheduled
20m
Aalborg University & Online

Aalborg University & Online

Presentation (with notebook) Application of AI in LCA F2 - Demo Derby

Speaker

Dr P. James Joyce (Watershed Technology Inc.)

Description

Highlights

  • PCFBench is a proposed benchmark for AI-supported Product Carbon Footprinting consisting of 6 independently scored subtasks, scoring systems on workflow stages rather than final kg CO2-eq values
  • Currently contains 614 expert-curated test items, from EPDs, practitioner-developed cases, and technical documents, released as open JSONL with a Python evaluation harness
  • Eight frontier LLMs performed poorly: claim-extraction F1 of 0.27-0.53; 25-55% of generated inventories failed mass-balance checks
  • Only 37-58% of final estimates fell within a factor of 2 of declared EPD results, demonstrating the need for task-level evaluation and domain-specific compound AI systems
  • We invite collaboration from LCA practitioners and developers to scrutinise, test and contribute to a shared reproducible evaluation infrastructure

Abstract

The use of AI systems in Product Carbon Footprint (PCF) modelling is becoming increasingly commonplace. Evaluation, however, is typically limited to cross-checks of final cradle-to-gate results or ad hoc tests of isolated capabilities using incompatible datasets and metrics. This makes it difficult to compare systems, identify where errors enter a workflow, or determine whether plausible results rest on defensible inventories and source evidence as opposed to merely reflecting category-typical values.

PCFBench is an open diagnostic benchmark for AI-assisted PCF modelling, covering six independently scored tasks: product decomposition; map-or-further-decompose triage; mapping of inventory items to secondary database activities; extraction of material input quantities (from technical documents); extraction of energy input quantities; and composition of these outputs into cradle-to-gate estimates.

Its 614 expert-curated items combine EPD-derived records, practitioner-authored cases, and evidence-grounded claims from technical literature. Each task has a typed input–output schema and task-appropriate metrics, allowing LLMs or other AI systems to be evaluated through common interfaces.

Across eight frontier LLMs, no model performed best on every task, and overall results were poor. Claim extraction from technical documents was consistently weak, with F1 scores of 0.27–0.53. In one cross-task analysis, 25–55% of generated life cycle inventories failed mass-balance checks, and only 37–58% of stepwise-composed final estimates fell within a factor of two of declared EPD results. PCFBench thus diagnoses where modelling errors arise, and shows that frontier LLMs alone are not yet fit for purpose for complex domain-specific tasks such as PCF.

In the spirit of BrightCon, we will present the open JSONL datasets, task definitions, schemas, scoring methods, Python evaluation harness, and observed failure modes. We invite the LCA community to scrutinise task boundaries and claims, test additional systems, and contribute data or task extensions, to help build shared, reproducible evaluation infrastructure for AI-assisted LCA.

How much time do you ideally wish for your contribution? 15 min (Presentation, slides; Presentation, with notebook)

Authors

Krishna Rao (Watershed Technology Inc.) Dr P. James Joyce (Watershed Technology Inc.)

Co-authors

Andrew Dumit (Watershed Technology Inc.) Daniel Frank (Watershed Technology Inc.) Gizem Ilayda Dinc (Watershed Technology Inc.) Jacob Feintzeig (Watershed Technology Inc.) Jonathan Glidden (Watershed Technology Inc.) Shaena Ulissi (Watershed Technology Inc.) Steven Watson (Watershed Technology Inc.) Travis M. Kwee (Watershed Technology Inc.)

Presentation materials