How we derive nutrition

Where the data comes from

Nutrient values in Canoli come from four published government food composition databases: CoFID 2021 (UK), Swedish Livsmedelsdatabasen 2026, Norwegian Matvaretabellen 2026, and the Canadian Nutrient File. Each ingredient carries its source name and external reference, retrievable via API and citable in published work. USDA is excluded by design.

CoFID 2021

UK Food Standards Agency / Public Health England. Open Government Licence v3.0. ~2,887 foods.

Swedish Livsmedelsdatabasen 2026

Livsmedelsverket (Swedish Food Agency). Open data terms. ~2,575 foods.

Norwegian Matvaretabellen 2026

Mattilsynet (Norwegian Food Safety Authority). CC BY 4.0. ~2,121 foods.

Canadian Nutrient File

Health Canada. Open Government Licence (Canada). ~5,690 foods.

Staged but not yet promoted: Frida v5.5 (DTU, Denmark, CC BY 4.0). Legacy McCance & Widdowson rows are present where not yet superseded by the current CoFID equivalent. Canoli-curated entries fill foods that are absent from any government source.

We exclude USDA. American mandatory fortification of grains (folic acid, iron, thiamin, riboflavin, niacin) and the broader American food universe make USDA values misrepresent the UK food supply for our users; we restrict to UK / EU / EEA / Commonwealth databases.

See the full source register →

The unit we work in

All nutrient values are reported per 100 g of edible portion, the UK FIC convention. Where the source database reports sodium only, salt is derived using salt = sodium × 2.5 / 1000. Sodium is stored in milligrams per 100 g; salt is shown to users in grams. Display rounding is value-dependent (see Rounding).

Salt vs sodium. Salt is derived from sodium where the source reports sodium only:

salt (g) = sodium (mg) × 2.5 / 1000

Where the source reports both, we cross-validate; disagreements greater than 10% are flagged for manual review, not silently corrected.

Rounding

Stored values keep full precision; rounding is applied only for display. Customer-facing labels and exports use the value-dependent EU rounding rules retained in UK law (Regulation (EU) No 1169/2011 and the European Commission's nutrition-labelling tolerance and rounding guidance). Each value is rounded to its tier first, then tested against the below-threshold rule.

Rounding is half-up, so a value sitting exactly on a boundary rounds up: 2.15 g is shown as 2.2 g.

Energy

Energy is computed using EU 1169/2011 Annex XIV conversion factors, applied uniformly across every source: protein 4 kcal/g, fat 9 kcal/g, available carbohydrate 4 kcal/g, fibre 2 kcal/g, alcohol 7 kcal/g. kJ is calculated independently of kcal rather than by 4.184 conversion, to prevent drift accumulating across complex recipes.

The factors:

This used to be source-aware: CoFID and legacy McCance carried a 3.75 kcal/g (16 kJ/g) factor because their carbohydrate column was expressed as monosaccharide equivalents (Greenfield & Southgate UK convention), and 3.75 compensated for the ~10% inflation that the mono-eq form produces in starchy foods. As of May 2026 we convert CoFID carbohydrate to by-weight at ingestion (see Carbohydrate), so the 4 kcal/g factor now applies everywhere.

kJ is calculated independently of kcal (not by 4.184 conversion) because the underlying factors differ by 1–2 kJ per 100 g per macronutrient and silently converting one to the other accumulates drift across complex recipes.

Recipe energy is derived, not summed. A recipe's energy is calculated from the recipe's combined macronutrients using the factors above, rather than by adding up each ingredient's stated energy figure. The two agree when every row is internally consistent; where they differ, the derivation from macronutrients is authoritative and the difference is disclosed on the recipe's derivation record. The one exception is an ingredient you supplied yourself with its own declared energy, which is honoured as declared (see Your own ingredients).

A disclosed divergence from Annex XIV. The Annex also assigns factors to organic acids (3 kcal/g) and polyols (2.4 kcal/g). We do not currently model either term, because source databases rarely report those masses. The effect is a systematic under-count of a few kcal per 100 g on moderately acidic foods (citrus and citrus juices, rhubarb, fermented products), usually within labelling rounding. Foods whose energy is dominated by organic acids, such as vinegars, are routed to nutritionist review rather than being given a derived figure. This is the only stated divergence from Annex XIV.

Source-reported energy is preserved alongside the recalculation. Differences greater than 10% flag the row for review.

Protein

Protein is stored in the regulatory form: total nitrogen × 6.25 (EU 1169/2011 Annex I), regardless of food type. CoFID rows are rescaled at ingestion from their food-specific Kjeldahl factors (5.18 for almonds up to 6.38 for dairy) back through to 6.25, food code by food code. Swedish, Norwegian, and Canadian sources already use 6.25 and pass through unchanged.

EU 1169/2011 Annex I defines protein for nutrition declarations as total nitrogen × 6.25, a single conversion factor regardless of the food. UK research databases historically use food-specific Kjeldahl factors instead (5.70 for wheat, 5.95 for rice, 6.38 for dairy, 5.71 for soya, 5.18 for almonds, 5.46 for peanuts and brazil nuts, 5.30 for other nuts and seeds, 5.83 for oats / barley / rye) to better reflect the amino-acid profile of the protein in that particular food.

Both numbers describe the same protein; they differ because the conversion factor differs. The research-form value is more accurate for nutrition science; the regulatory form is what an EU food label must declare.

Canoli stores the regulatory form. For CoFID rows we rescale the published value back through the food-specific factor and forward through 6.25, food code by food code. For wheat-based products this raises protein by approximately 9.6% (factor 5.70 → 6.25); for dairy it lowers protein by approximately 2.0% (factor 6.38 → 6.25); for almonds it raises it by 20.6% (factor 5.18 → 6.25). Sources that already use 6.25 (Swedish, Norwegian, Canadian) pass through unchanged.

Carbohydrate

Carbohydrate is stored by weight (the EU 1169 requirement), not as the monosaccharide equivalents CoFID publishes. CoFID rows are converted to by-weight on ingestion using the row's own starch and sugar breakdown, typically about 9% lower than the historic mono-equivalent value. Swedish, Norwegian, and Canadian sources report by weight directly.

EU 1169 requires available carbohydrate (mono- + disaccharides + starch and other polysaccharides excluding fibre), reported by weight. Sources differ on how they express this:

In that conversion, total carbohydrate, the starch line, the total sugars line ("of which sugars" on the label), and the available carbohydrate line are all rewritten:

starch_by_weight = starch / 1.10

sugars_by_weight = (glucose + fructose + galactose)
                 + (sucrose + maltose + lactose) / 1.05

carbohydrate_by_weight = starch_by_weight + sugars_by_weight

Where the per-sugar breakdown is incoherent with the total sugars line (sub-columns sum more than 0.2 g or 10% away), we fall back to total sugars / 1.05 for both the sugars line and its contribution to carbohydrate. Where only one of starch / sugars is reported, we derive the missing fraction from the row's total carbohydrate. Where neither is reported (rare; some dried mushrooms and herbs), the row is left unconverted and flagged for nutritionist review.

For a typical wheat product the by-weight carbohydrate is approximately 9% lower than the historic monosaccharide-equivalent value.

Whether the figure includes fibre is also recorded. European sources report available carbohydrate (fibre excluded). North American sources, including the Canadian Nutrient File, report carbohydrate by difference, which includes fibre. We record which convention each row uses and honours it in every downstream calculation: for a by-difference row, fibre is deducted when deriving available carbohydrate and energy, so fibre is never counted twice. The stored figure itself is never altered to fit our reading of it.

Where fibre is missing, we default it to 0 and flag the row with an assumed-zero indicator. Free sugars are a distinct sub-component, handled separately under Free sugars below.

Vitamins and equivalents

Vitamin A is normalised to Retinol Activity Equivalents (RAE) per IOM 2001 / EFSA 2015: 1 µg RAE = 1 µg retinol = 12 µg β-carotene. Vitamin E sums tocopherol fractions where reported. Niacin Equivalents, Dietary Folate Equivalents, and α-Tocopherol Equivalents are not computed, so %NRV from these vitamins reads as source-reported rather than bioavailability-adjusted.

Vitamin A is normalised to RAE (Retinol Activity Equivalents per IOM 2001 / EFSA 2015):

1 µg RAE = 1 µg retinol = 12 µg β-carotene = 24 µg α-carotene / β-cryptoxanthin

Vitamin E. Where sources provide tocopherol fractions (Canadian: α / β / γ / δ), they are summed into vitamin_e_mg. Other sources report a single value.

Folate. Stored as total folate (µg). The user-facing label is "Folic acid" per FIC NRV vocabulary: the regulation uses the synthetic-form name for the NRV, and we mirror the regulation rather than the chemistry.

Cross-source reconciliation

Each ingredient is matched across sources by a layered strategy: exact name match, fuzzy text similarity, semantic meaning-level matching, and manual review for borderline cases. The composite ingredient inherits its primary value from the highest-priority source for the user's context (CoFID first for UK users); other sources gap-fill missing nutrient cells.

Each source ingredient is normalised: lowercased, processing state extracted (raw / boiled / fried / etc.), form extracted (lean only, kernel only, etc.). Matching strategies are layered:

  1. Exact match on normalised name.
  2. Fuzzy text similarity for spelling and word-order variation.
  3. Semantic matching for meaning-level equivalence.
  4. Manual review for borderline cases below the auto-match threshold.

Matched ingredients form a composite record. The composite inherits its primary value from the highest-priority source for the user's context (CoFID first for UK users); other sources gap-fill missing nutrient cells.

Provenance is preserved: every gap-filled value records which source supplied it, retrievable per ingredient via the API. Where sources disagree beyond tolerance, the composite uses the primary value and surfaces the disagreement rather than averaging them.

Quality gates

Every source ingredient passes five gates at ingestion: identity (source, reference, name must be present), FIC-7 mandatory (at least 4 of the 7 mandatory nutrients), biological range, internal consistency, and deduplication on (source, external reference). Further validation adds macro closure, free-sugar completeness, and energy plausibility.

Identity

Source + external reference + name must be present, else rejected.

FIC-7 mandatory

At least 4 of {energy, fat, saturated fat, carbohydrate, sugars, protein, salt or sodium}, else rejected. All 7 → "complete"; 4–6 → "partial".

Biological range

Each value bounded (energy 0–900 kcal, fat 0–100 g, sodium 0–40,000 mg, etc.). Out-of-range values are flagged, not rejected.

Internal consistency

Saturated ≤ total fat; sugars ≤ carbohydrates; computed energy within 10% of source-reported.

Deduplication. Records are unique on (source, external reference); a newer version of the same record updates in place rather than duplicating.

Further validation additionally covers negatives, macro closure, name/nutrient mismatch, per-100 ml detection (liquids accidentally reported on volume basis), free-sugar completeness, and energy plausibility.

The same internal-consistency checks run continuously against the live catalogue and against ingredients customers supply themselves; a row whose stated energy stops agreeing with its own macronutrients is flagged rather than silently corrected.

Of approximately 13,475 ingredient records currently held, 0 are flagged for manual review (cited as evidence the gates work and matter).

Recipe-level calculation

Recipe per-100 g is the weighted average of each ingredient's per-100 g, weighted by mass fraction of total raw mass. Yield converts raw mass to finished mass: yield below 100% concentrates nutrients; yield above 100% dilutes them (pasta, rice, lentils). Fruit / veg / nut / seed percentage is computed for HFSS scoring with the 2018 dried-fruit weighting.

Yield (recipe-level or per-ingredient %) converts raw mass to finished mass:

Yield is clamped to ≥1% to prevent division collapse. Per-serving = per-100 g × (serving_grams / 100). Fruit / veg / nut / seed percentage (FVNS%) is computed for HFSS scoring with the dried-fruit weighting prescribed by the 2018 model.

One basis, shared by every reported output. A recipe carries a single as-sold nutrition result. The on-screen nutrition panel, exported labels and documents, nutrition claims, and the NPM/HFSS and Nutri-Score assessments are all calculated from that same result.

Cooking retention is applied per ingredient line, and that is a deliberate departure. The published retention factor tables (Bognar 2002; EuroFIR) are dish-level: one factor applied to the finished dish. We apply each ingredient line's own factor instead. On the one dish in the source's own validation set where both approaches can be tested against analysed values, per-line application gives the smaller error (about 3% on fat, against about 15% dish-level), which is why we departed from the published procedure. Where a line's method or nutrient has no published factor, the raw value is kept and the line is flagged rather than estimated.

Your own ingredients

Ingredients you add yourself, by hand or from a scanned specification, carry your declared values, and your declared values are authoritative: we never overwrite them with our own derivation. Consistency checks run on entry and show you any disagreement between your figures; they warn, they do not correct.

Many of our customers work from supplier specifications and have gone to print with those figures. So the rule is absolute: a value you declared is a value we keep. If your declared energy differs from what EU 1169 Annex XIV derives from your declared macronutrients, your recipes' nutrition, labels and scores are calculated with your figure, and the difference is disclosed on the recipe's derivation record rather than silently adjusted.

On entry, and whenever a value changes, the same internal-consistency checks that police our own catalogue run on your ingredient and any findings are shown to you:

A warning stays visible until the underlying values change. Nothing is auto-corrected, and a warning never blocks you from using the ingredient. Where your supplier's specification uses a different carbohydrate convention (see Carbohydrate), the convention is recorded and honoured rather than the figure being altered.

Free sugars

Free sugars follow the SACN 2015 definition: added mono- and disaccharides plus those naturally present in honey, syrups, fruit juices, and purées. Intrinsic sugars in whole or cellularly-intact fruit and vegetables are excluded. Where the source provides a free-sugars value (Swedish), Canoli uses it directly; otherwise rule-based classification flags each food by name pattern, with a nutritionist reviewing borderline cases.

Definition per SACN 2015: added mono- and disaccharides plus those naturally present in honey, syrups, fruit juices, and purées. Intrinsic sugars in whole or cellularly-intact fruit and vegetables are excluded.

Logic:

Classification rules and every unclassified → classified transition are reviewed by a registered nutritionist before being promoted to production.

Claims, NPM and HFSS

Fifteen EU 1924/2006 nutrition claims are evaluated against per-100 g thresholds for energy, fat, saturated fat, sugar, sodium, fibre, and protein. "Source of" and "high in" vitamins and minerals are evaluated separately at 15% and 30% NRV per 100 g. Both UK 2004/05 and 2018 NPM/HFSS models are implemented, and both are validated against every published worked example: the 2018 model against the seventeen examples published with it, and the 2004/05 model against the six examples in the 2011 technical guidance.

15 nutrition claim types per Reg (EC) No 1924/2006 Annex evaluated:

"Source of" and "high in" vitamin or mineral: 15% / 30% of NRV per 100 g, against EU 1169/2011 Annex XIII Part A.

The engine flags qualifying claims; it does not assert the legality of any specific marketing statement, which depends on conditions of use (claim wording, comparator product, target population) outside the threshold check.

NPM and HFSS. Both UK 2004/05 and 2018 Department of Health Nutrient Profiling Models are implemented, and both are validated against every worked example their guidance publishes. The 2018 model passes all seventeen worked examples published alongside it. The 2004/05 model, which is the version currently in force for HFSS promotion and advertising, passes all six worked examples in the 2011 technical guidance. These validations run in our test suite on every change.

The 2011 guidance publishes two fibre threshold scales for the 2004/05 model, one for NSP (Englyst) values and one for AOAC values, and directs that the NSP scale be used where an NSP value is known. Canoli currently scores all fibre against the AOAC scale. We are recording each value's analytical basis through the ingredient pipeline so that the matching scale can be selected per value; until that lands, a score derived from an NSP measurement is scored on the AOAC scale, which under-credits fibre. Where only an NSP value is available and an AOAC basis is needed, a factor of 1.33 (Lunn & Buttriss, 2007) is applied and the value is disclosed as an estimate; that factor is a population-level equivalence rather than a per-food conversion.

The protein-points exclusion rule (A ≥ 11 and FVN < 5 → protein zeroed) is applied per the original DH guidance.

Nutri-Score. The recipe builder can also show a Nutri-Score A to E grade, computed with the 2023 Santé publique France algorithm (the 2022 main-algorithm and 2023 beverages updates published by the Nutri-Score Scientific Committee) from the same per-100 g values. Its scoring is validated against the official published Nutri-Score reference cases (Santé publique France, 2023). It is a voluntary front-of-pack scheme used in parts of the EU, offered here as an at-a-glance comparison aid; it is not a UK regulatory label and does not replace the NPM/HFSS verdict, which remains the UK regulatory basis. An organisation can switch the panel off in its settings, and a recipe with an unknown required nutrient is shown as not yet assessable rather than being given a grade.

Fibre methodology citations. Where a source database publishes only NSP fibre values, Canoli estimates the AOAC value as NSP × 1.33 (Lunn & Buttriss, 2007) so the displayed value can sit alongside AOAC-native values without methodology drift. This factor is a population-level equivalence rather than a per-food conversion, so estimated values are flagged as such. The references that establish and corroborate this approach:

Allergens

Canoli classifies the 14 UK mandatory allergens per FIC Annex II using word-boundary pattern matching against curated keyword sets, with named exclusions for common false positives (almond milk does not fire the milk rule; coconut does not fire the tree-nut rule). Composite recipes inherit allergens from components. Ambiguous cases stay flagged for nutritionist review rather than silently asserting allergen-free.

Canoli classifies the UK 14 mandatory allergens per FIC Annex II: cereals containing gluten, crustaceans, eggs, fish, peanuts, soybeans, milk, tree nuts, celery, mustard, sesame seeds, sulphur dioxide and sulphites at >10 mg/kg, lupin, molluscs.

Classification combines:

All inferred classifications are tagged with confidence "inferred"; promotion to production allergen status requires nutritionist review. The system never silently asserts allergen-free for ingredients it cannot classify with high confidence; ambiguous cases stay flagged rather than guess.

Food categories

Every ingredient is filed into a single food category for grouping and browsing. Where a source's own grouping is missing or misleading, we apply the UK/EU regulatory definition rather than the source label.

For meat we follow the WHO definition of processed meat: meat transformed by salting, curing, fermentation, smoking, or the addition of ingredients to enhance flavour or improve preservation. The two buckets are:

Freezing on its own does not make a food processed; the curing or the added ingredients do. Mechanically separated meat is filed as processed.

The derivation record

Every recipe can export a derivation record: a document stating each reported value, the source behind each ingredient line, the assumptions in force, and any open methodological questions that touch the recipe. Publishing a recipe version freezes its numbers as computed at that moment, and a restored older version is treated as a new publication.

The derivation record is the document to file alongside a label or a specification. It shows, per value, where the number came from and what was assumed in producing it, including any disclosed difference between a declared figure and its derivation.

Published numbers do not drift. When you publish a recipe version, its results are frozen exactly as computed. Later changes to source data, to our method, or to the recipe itself never alter a published version's numbers; they apply to the next version you publish. Restoring an older version is itself a publication, so the restored numbers are frozen afresh rather than inherited.

Open questions are registered, not silently decided. Where the correct treatment is genuinely unsettled, for example the energy factor for rare sugars such as allulose, which UK labelling law does not define, the question is logged on a rulings register, the assumption currently in force is stated in plain language on any affected derivation record, and the eventual resolution is applied to future calculations rather than rewriting past ones.

What this page doesn't claim

Bioavailability is not modelled in %NRV: iron from spinach and iron from beef display identically. Cooking nutrient retention, where a cooking method is set, is applied to the recipe's single as-sold result, which every output shares. Fortification is not separately flagged. Niacin Equivalents, Dietary Folate Equivalents, and α-Tocopherol Equivalents are not computed. Every nutrient carries an explicit state: Known (a measured or derived value, including a genuine zero), Unknown (no value from any source, shown as "Unknown" rather than zero), Trace (present below the level it can be quantified), or Not applicable. A missing value is treated as Unknown, never silently as zero.

How to verify, how to push back

Every ingredient in the Canoli API exposes source and external_ref fields suitable for academic, regulatory, or label-substantiation citation. The composite-ingredient endpoint returns gap-fill provenance per nutrient, so a citation can be made specific to which source supplied which value. Corrections go to data@canoli.co.uk and are reviewed by a registered nutritionist.

A citation built from those fields looks like this:

Salmon, raw - Canoli ingredient ID 4123
Primary source: CoFID 2021 (UK FSA, Open Government Licence v3.0)
Source ref: food code 11-678
Vitamin D: 8 µg/100g (gap-filled from Norwegian Matvaretabellen 2026, food code 0445)
Retrieved: 2026-04-15