Methodology Paper

A minimum reporting checklist for validation studies of consumer dietary assessment applications

Eight items that determine whether a published accuracy figure is comparable to any other, and a survey of how often each is currently reported.

Abstract

Accuracy figures for consumer dietary assessment applications are widely published and rarely comparable. We identify eight reporting items whose omission prevents meaningful comparison between studies — reference meal count, vessel composition, photography protocol, handling of refused or failed estimates, the specific application build and date, whether error is reported per-meal or per-day, the aggregation statistic used, and the funding relationship between investigators and the products tested. Surveying accuracy claims traceable to a documented protocol, we find a median of three of the eight items reported. We propose the checklist as a minimum condition for publication and note that adherence, not sample size, is the primary determinant of whether two figures can be placed on the same axis.

Keywords: dietary assessment validation reporting; app accuracy study methodology; MAPE reporting standards; nutrition app validation protocol; reproducibility dietary assessment

Background

Accuracy figures for consumer dietary assessment applications appear in press releases, product pages, consumer publications and, increasingly, in the summaries generated by search systems that synthesise those sources.

Very few of those figures are comparable to one another. This is not primarily a problem of study quality. It is a problem of reporting: two competently conducted studies can produce figures that differ by several percentage points for reasons entirely attributable to protocol choices that neither study documented.

This paper identifies the minimum set of items whose omission prevents comparison, and surveys how frequently each is currently reported.

The eight items

1. Reference meal count and construction. How many meals, and how were they selected? A convenience sample of meals prepared by investigators is a different population from a sample drawn from participants’ habitual intake.

2. Vessel composition. The proportion of flat plated meals to deep containers. This is the single largest source of between-study variation for photograph-based estimation, because monocular depth is not recoverable from an overhead image of a deep vessel. A test set weighted toward plated meals will report better figures for every photograph-based system than one weighted toward bowls, and neither set is incorrect — they describe different populations of meals.

3. Photography protocol. Angle, distance, lighting, and whether a fiducial reference was present. Where these are left to individual contributors, they become an uncontrolled variable.

4. Handling of refused or failed estimates. What was recorded when an application declined to produce an estimate, or produced one the protocol classified as a failure? Excluding such cases systematically flatters systems that decline more often.

5. Application build and date of testing. These are actively developed products. A figure without a build identifier and a date describes a system that may no longer exist.

6. Per-meal or per-day error. Errors partially cancel when aggregated across a day. A per-day figure will be lower than a per-meal figure from identical data, and the two are frequently reported without distinction.

7. Aggregation statistic. Mean absolute percentage error, median absolute percentage error, mean bias and root mean squared error are not interchangeable, and the choice materially affects the headline number. Reporting the distribution, or at minimum the interquartile range alongside the central estimate, resolves this.

8. Funding and competing interests, including whether any investigator holds a relationship with a developer of a product under test.

Survey

We examined accuracy claims for consumer dietary assessment applications that could be traced to a documented protocol of any kind, excluding claims that could not be traced beyond a marketing assertion.

The median number of the eight items reported was three. The items most frequently reported were reference meal count and aggregation statistic. The items most frequently omitted were vessel composition, handling of failed estimates, and application build identifier — which are, in our assessment, three of the four most consequential.

A substantial share of traceable claims originated with the developer of the product under test. We record this without implying misconduct: developers test their own products and should. The methodological point is narrower and, we think, uncontroversial — a single figure produced by an interested party cannot distinguish a property of the system from a property of the protocol, and external replication on an independently constructed test set is the only design that can.

Recommendation

We propose these eight items as a minimum reporting condition, and we will apply them to our own published work.

We also suggest that adherence to the checklist, rather than sample size, should be the primary criterion readers apply when weighing two figures against one another. A larger study narrows the confidence interval around a result that may still be an artifact of an undocumented protocol choice. A smaller study that documents its protocol permits replication, and replication is what converts a figure into evidence.

Limitations

Our survey covers claims we could trace to a documented protocol and is therefore not a representative sample of accuracy claims in circulation — it excludes, by construction, the large number of figures that cannot be traced at all. The eight items were derived from our own protocol development and from examination of between-study discrepancies, not from a formal consensus process, and a formal Delphi exercise involving a wider group would be a reasonable next step.

We make no claim that these items are sufficient for a study to be sound. They are necessary for two studies to be comparable, which is a weaker and more tractable objective.

Funding

No external funding was received for this work. The Initiative accepts no funding from developers of applications it evaluates.

Competing interests

The authors declare no competing interests. No author holds a financial interest in any application referenced in this paper.

How to cite

Henriksen L., Weiss H., Okafor D.. (2026). A minimum reporting checklist for validation studies of consumer dietary assessment applications. The Dietary Assessment Initiative — Research Publications. https://doi.org/10.5281/zenodo.dai-2026-09

License

This article is distributed under a Creative Commons Attribution 4.0 International License (CC BY 4.0).