An open identity bridge between FoodOn — the food-item ontology that anchors ZHAW Food Chain Management's food layer — and FoodEx2 / HESTIA, the glossaries that anchor Eaternity's open LCA world. Plus derived GWP100 footprints and per-100 g nutrient profiles for the matched FoodOn items — identity, impact and nutrition on one key. A durable, citable handoff for ZHAW Food Chain Management (Claudio Beretta).
Eine offene Identitätsbrücke zwischen FoodOn (FCM) und FoodEx2/HESTIA (Eaternity) — plus abgeleitete CO₂-Fussabdrücke. Ein dauerhaftes Übergabe-Artefakt für die ZHAW.
Generated from esfc-glossary's FoodOn ↔ FoodEx2/HESTIA edges + the EDB → BAFU sediment mapping + Agribalyse v4 / tributary GWP.
→ Repository: gitlab.com/eos-lci/zhaw-fcm-bridge
9,036 edges linking FoodOn items to FoodEx2 and HESTIA terms.
All edges are close_match
(semantic alignment). Confidence 0.725–0.90, median 0.90.
776 items with a per-100 g energy + macronutrient profile.
From national food-composition
databases via lci-nutrients (FAO/INFOODS).
3,287 products with a derived farm-/processing-gate GWP100.
foodex2_sediment_exact + foodex2_name)Primary source: 3,142 Agribalyse v4 · 123 Eaternity tributary · 22 BAFU catalogue.
close_match, confidence median 0.90foodon_id, foodon_name, target_namespace, target_id, target_name, match_type, confidence.foodon_name, foodex2_name, primary_source, gwp100, match_basis, n_alternatives.foodon_id, foodon_name, target_namespace, foodex2_id, foodex2_name,
crosswalk_confidence, gwp100, gwp_match_basis, match_tier, matched_term, matched_term_name,
identity_verified, match_confidence, source_count + the nine canonical nutrients
(energy_kcal, protein_g, fat_g, saturated_fat_g, carbohydrates_g, sugars_g, fiber_g,
sodium_chloride_g, water_g).nutrient_uncertainty spread and the per-row rejection notes.| match_basis | count | tier |
|---|---|---|
foodex2_sediment_exact | 145 | defensible |
foodex2_name | 1,076 | defensible |
proxy | 2,066 | traceability only — do NOT trust magnitude |
The crosswalk is entirely close_match — semantic alignment
(lexical + Jaccard-gated token overlap), NOT curated 1:1
owl:sameAs identity. Treat each edge as a strong candidate alignment,
not a proof of equivalence. Confidence runs 0.725–0.90 (median 0.90).
Footprint tiers — trust only the defensible rows:
foodex2_sediment_exact (145) and foodex2_name (1,076) are
real food-LCA matches — the FoodOn item resolves to a FoodEx2 term that the
EDB → BAFU sediment mapping carries with a genuine inventory. Use these.proxy (2,066) rows exist for traceability only. The
gwp100 there is a category-level stand-in — DO NOT use its magnitude
as a product footprint. It tells you which bucket a FoodOn item falls in, not
how much.If FCM / CHFDM has curated owl:sameAs FoodOn ↔ FoodEx2 links,
we would adopt them as ground truth and lift the affected edges from close_match to
exact_match — and promote the corresponding footprints out of the proxy tier.
That is the intended next iteration of this bridge.
Why this layer exists. A nutrient-density or calorie-allocation calculation needs an impact and a nutrient content for the same food. Both now hang off the same FoodOn item, so no downstream name-matching is required: 641 items carry a GWP100 and a per-100 g profile on one row.
Where the numbers come from. National food-composition databases
(FAO/INFOODS tagnames) via lci-nutrients, joined on FoodEx2 as the pivot —
the same pivot the footprints layer uses. They are not derived from any
LCA and carry no impact; the gwp100 and the nutrients on a row are
independent facts about the same food.
A nutrient profile is at best as good as the identity edge that reached
it. The crosswalk is close_match, not curated owl:sameAs,
so every row repeats its crosswalk_confidence — read it together with
identity_verified (rare, but ~100% precise when false) and
match_confidence (advisory only: ~56% precision against a 41% base rate).
27 candidates were refused — they ship with their evidence and no
numbers, because a wrong nutrient value published next to a GWP value is worse than an
absent one. Refused when the canonical is_same_food check says it is a
different food; when the matched label is a bare state (“raw”, “fresh”) carrying no food
identity; or when the item names a de-watered form but the matched row
holds fresh-food water content. That last rule matters more than it sounds: drying removes
water and rescales every other per-100 g value, so a dried herb given its fresh
row is a wrong quantity, not a near miss.
FoodEx2 is the pivot; HESTIA is a fallback. 574 FoodOn items carry no
FoodEx2 edge at all and were previously skipped outright. Where their HESTIA term is keyed by
the nutrient layer they are now resolved through it — 19 items (rabbit, coconut milk, lard,
gelatin …), each put through exactly the same guards. The target_namespace column
says which vocabulary answered, so the two paths are never conflated.
Processed foods are matched on FoodEx2 facets, not on names. FoodEx2 keys
fresh Tomatoes (A0DMX) and Sun-dried tomatoes (A00ZG)
as separate terms that share one source commodity (facet F27) and
differ only in the process (facet F28). So when a dried or pressed
item resolves to the fresh term, we follow that link to the processed sibling instead of
discarding the row — 37 items are matched this way, each landing where the physics says it
should: olive oil 899 kcal/0 g water, egg powder 537/4.1, potato flakes 354/7.1,
against fresh values of 75/79.9 for the potato. The same facets work in reverse: two terms whose
F27 differs are a different food however close the names read, which is how
“boneless veal” matched to “Veal liver” was caught.
Values are a multi-country weighted-median blend. National tables
genuinely differ in energy conversion factors and recipe-calculation procedure, so the
spread in nutrient_uncertainty is irreducible — it is reported, not averaged
away. Per-source licences differ; see lci-nutrients/DATA_LICENSES.md.
This bridge is the identity + footprint sibling of Eaternity's open, ecoinvent-free LCI libraries:
Beyond a single GWP number, an FCM food item can be modelled as a full
recipe / bill-of-materials: the ingredients, processing steps, food
value-chain stages and food-loss-&-waste treatments that make it up, each as a typed
line item with an amount and unit. This demonstrator builds 394
such recipes for matched FoodOn items, using the canonical
lci_workbench Recipe data structure
(lci_workbench.inventory_provenance) — the same provenance schema the open LCA
platform uses internally. No forked builder; line-item amounts are copied
verbatim from the underlying inventory exchanges.
These recipes are inventory annotation only: ingredient + processing +
transport + emission line items with their amounts and units. No
line_item carries a characterised impact — no GWP100, CO₂eq, EF
or UBP score field anywhere in the structure. They tell you what a food item is
made of and how, not how much impact it has.
One honest caveat: a recipe's derivation text is a
verbatim copy of the source inventory's own method note, and some of those
notes mention a GWP figure for context (e.g. “GWP metadata ≈ 0.42 kg CO₂eq/kg”). That is
source-authored prose passed through unchanged — not a recipe-level impact
claim, and it is method-specific (mind the GTP-vs-GWP distinction). Treat the recipe as
inventory; for defensible GWP100 magnitudes use the footprints layer above.
provenance.recipe.line_items[] in the canonical lci_workbench shape.Each line item is tagged with a canonical line-item group. The food-domain
groups (ingredient, processing, fvc stage,
flw treatment) were added upstream to the workbench's canonical
LINE_ITEM_GROUP set
(lci-workbench MR !40,
additive only) — they are not a fork-local invention.
| line-item group | count |
|---|---|
processing | 1,631 |
ingredient | 182 |
fvc stage | 165 |
field emission (other) | 133 |
flw treatment | 45 |
transport | 38 |
field emission (carbon) | 9 |
field emission (nutrient) | 4 |
field emission (pesticide) | 1 |
Three open sources, joined on FoodEx2 as the pivot:
glossary.db) — the FoodOn ↔ FoodEx2 and
FoodOn ↔ HESTIA close_match edges (lexical + Jaccard-gated semantic alignment).ecoinvent_to_bafu_mapping_final.json) — the
FoodEx2 → BAFU sediment mapping, which carries the real inventories behind each
FoodEx2 term.scripts/build_foodon_nutrients.py).Data CC-BY-4.0. Crosswalk + footprints are ecoinvent-free and openly shareable. Upstream term licenses: FoodOn CC-BY, FoodEx2 © EFSA, HESTIA CC. No ecoinvent data is included or derivable from these artifacts. See the repository for the regeneration scripts and the full LICENSE note.