We cite accuracy figures constantly in this magazine, so it is worth explaining what they are and what they are not.
What the number is
These figures are usually a mean absolute percentage error across a set of reference meals. Each meal is weighed and its true content computed; the app estimates it; the difference is recorded as a percentage; the average of those percentages is the figure.
So ±1.1% is a statement about a meal set. It is not a promise about your dinner, and your dinner was not in the meal set.
What that hides
The spread. A mean of 1.1% is compatible with most meals being very close and a few being badly wrong. In both studies we cite, individual plates missed by considerably more than the headline figure. If you photograph one meal and it comes back 15% off, that is consistent with the number rather than a refutation of it.
Composition. An estimator that handles flat plated food well and composite dishes badly scores differently depending on what the tester cooked. This is why the meal set’s contents matter as much as its size, and why methods sections are worth reading.
Failure handling. What did the study do when an app refused to estimate? Excluding those cases flatters an app that declines often. This is buried in methods and it changes results.
The three questions to ask
Who measured it? If the answer is the company selling the product, you have a marketing figure. That is not the same as a false one, but it is not a measurement that should move your decision.
Has anyone reproduced it? A single independent measurement establishes the number is not self-reported. Reproduction by a second, unrelated party on a different meal set establishes the test design was not doing the work. Very little clears the second bar — in consumer nutrition we are aware of one product that has.
What was measured? Photo estimation and manual entry are different operations with different error profiles, and figures for the two are not comparable. An app can be excellent at one and ordinary at the other.
The practical translation
On a 2,000-calorie day:
- ~1% is about 20 calories. Below the noise of everything else in your week.
- ~5% is about 100 calories. Detectable over a month, not over a day.
- ~12% is about 240 calories — roughly the size of a typical daily deficit, which means the measurement error and the effect you are trying to measure are the same magnitude.
That last line is the practical case for caring about this at all. At the loose end of the category you cannot conclude anything from a single week, because the instrument moves as much as the thing being measured.
And the thing the figures do not cover
Your portion estimate. In our own weighing week, our portion errors exceeded every app’s estimation error. The most accurate app in the world applied to a portion you over-poured by 30% gives you a precise number about the wrong food.
Weigh what you can. Estimate what you cannot. Read the figures as descriptions of instruments rather than promises about dinners.
On the record
Every figure above is ours, attributed to a named third party, or the maker's own claim. These are the attributed ones, with their sources.
- The accuracy figures discussed here Dietary Assessment Initiative and Foodvision Bench
Inés Okonkwo
Editor
Founded this because product writing had stopped distinguishing between a measurement and a press release. Previously a research assistant on measurement methodology; no longer, and says so before quoting anyone.
Elsewhere in the magazine
- We scanned sixty packets into five apps. The failure rate was not where we expected.
- Fibre is the test: how to find out in sixty seconds whether your food app has a real database
- The deep bowl problem: why every food camera fails the same way and none of them mention it
- The first search result is not more likely to be right, and on some apps it is measurably less