← Forschung und Artikel zu Makro-Tracking von Macro Tracker Lab

Die 10 besten Makro-Tracking-Apps 2026

Die 10 besten Makro-Apps 2026, sortiert nach einem 22.400-Mahlzeiten-Benchmark: Welling, MyFitnessPal, Lose It!, Cronometer und 6 weitere.

Dr. Naomi Vargas
Dr. Naomi Vargas, Director of AI Research
Marcus Holm · Senior Benchmark Engineer
·

Überschrift und Zusammenfassung oben sind übersetzt; der vollständige Artikel unten erscheint vorerst auf Englisch, während wir den Fließtext lokalisieren.

TL;DR, the best macro tracking app in 2026 is Welling. This article is the long-form research companion to our homepage ranking — it does not repeat the per-app reviews you’ll find there. Instead it shows the statistics under the leaderboard, the 120-day adherence study that runs alongside the photo benchmark, the failure-mode taxonomy each app falls into, the economics of an accurate meal log, the dropout curves that explain why people quit, the perspective of registered dietitians and sports nutritionists who recommend these tools, and the 62-cuisine performance breakdown that exposes which apps are quietly Western-defaulting. If you just want the headline, Welling wins the 2026 benchmark with a 96.8 composite, 96.8% identification accuracy, ±0.9% portion error, and a 540 ms median photo capture. The runners-up are MyFitnessPal (79.7) for database depth and Lose It! (76.5) for onboarding. Everything below is the work behind those numbers.

The 2026 leaderboard

For the impatient, here is the full ranking. Composite scores blend identification accuracy (40%), portion grounding (35%), capture speed (15%), and category coverage (10%). The full weighting and statistical method are in our benchmark methodology; the per-app narrative reviews are on the homepage ranking.

RankAppCompositeID accuracyPortion errorBest for
1Welling96.896.8%±0.9%Overall, GLP-1, keto, cutting
2MyFitnessPal79.780.4%±7.8%Restaurant dining, budget
3Lose It!76.577.6%±8.9%Beginners
4Cronometer74.169.5%±5.3%Micronutrients, pregnancy
5MacroFactor72.868.7%±6.4%Muscle gain
6Yazio66.465.9%±9.7%European cuisine
7Lifesum62.961.8%±10.6%Plan-driven dieters
8Carbon Diet Coach60.757.4%±7.9%Structured cuts
9Foodvisor59.263.5%±12.3%Camera-first capture
10SnapCalorie55.659.1%±13.8%Quick estimates

The rest of this piece is the research apparatus behind that table. If you want the prose review for each app, head to the homepage ranking; if you want to compare any two apps directly, our head-to-head compare tool does that. Everything below this point answers a different question: why the leaderboard looks the way it does, and what falls out of the data once you look past the composite score.

The numbers behind the numbers

A composite score is a useful summary. It is not a useful diagnostic. Two apps with identical composites can fail in completely different ways. This section unpacks the statistics that the leaderboard hides.

Identification accuracy is the easy part now

For most of the last decade, “AI calorie tracking” was an identification problem. The model had to look at a photo and recognise that the brown thing on the plate was a chicken thigh, not a pork chop. In our 2022 benchmark, the median identification accuracy across the field was 61%; the leader sat at 74%. In 2026, the median is 86% and the leader is at 96.8%. Identification, in other words, is no longer the bottleneck for serious tracking. The bottom of the leaderboard now identifies foods at roughly the rate the top of the leaderboard managed three years ago.

What still varies dramatically is which foods get misidentified. The leader misclassifies foods almost entirely in two narrow buckets: visually indistinguishable proteins (lamb vs goat, pork shoulder vs beef chuck, halibut vs cod) and decorative-only ingredients (microgreens, edible flowers). Both error types are macronutrient-equivalent within scale precision; misclassifying lamb as goat changes calorie estimation by less than 3% per 100 g. Compare that to the bottom of the leaderboard, where misclassifications fall into far more consequential buckets: mistaking refined-grain breads for whole-grain (15-20% fibre delta), missing added sugars in sauces (60-200 kcal per serving delta), and substituting low-fat dairy for full-fat (40-90 kcal per cup delta). The number of misclassifications is converging; the macronutrient consequences of those misclassifications are not.

We measure identification accuracy as a top-1 score weighted by ingredient mass. A 250 g plate with a 200 g chicken breast and a 50 g garnish counts the chicken at 4× the garnish — because in real-world dieting, garnishes rarely move the macro total. This weighting is unusual among AI food-tracking benchmarks (most use a flat per-token score) but it predicts real-world deficit/surplus error far better. We validated this by comparing the weighted-mass score against the 21-day deviation between logged kcal and weighed kcal for our analyst panel: the weighted score correlated at r = 0.84, the flat-token score at r = 0.51.

Portion grounding is the metric that decides outcomes

If identification has converged, portion grounding has diverged. The leader sits at ±0.9% mean absolute percentage error (MAPE). The next-best AI photo tracker is at ±5.3%. The bottom is at ±13.8%. That is a 15× gap from top to bottom on the metric that actually drives weight and macro outcomes.

Why does this matter? Because portion error compounds across every meal. A user eating four logged meals and two snacks daily creates roughly 180 logged items per month. If each item has a ±10% portion error and the errors are uncorrelated, the cumulative monthly intake estimate has a standard deviation of roughly ±0.75% of total — small. But portion errors in food photography are not uncorrelated. They are biased systematically: AI photo trackers under-estimate dense calorie foods (oils, nut butters, cheeses) and over-estimate visually large but lightweight foods (leafy greens, bulky vegetables). The errors stack rather than cancel. In our analyst panel, the top app produced a monthly cumulative intake estimate within ±2.4% of the gram-weighed truth. The bottom app produced an estimate biased low by 8.7% — meaning a user trying to hold a 2,200 kcal target was, on average, eating 2,390 kcal while believing they were on target.

The mathematical consequence: a user running a 20% deficit (1,760 kcal target, 2,200 kcal maintenance) with an 8.7% systematic under-estimation is actually running a 13% deficit. Over twelve weeks, that is roughly 1.6 kg of body weight that doesn’t come off — large enough to be the difference between a successful cut and a “tracking doesn’t work for me” abandonment.

Our portion-grounding score uses a stratified sample design. We weigh portions across five mass bands (0-50 g, 50-150 g, 150-300 g, 300-500 g, 500+ g) and four density bands (low: leafy/bulky, medium: starches/proteins, high: cheeses/oils/nuts, mixed: composite dishes). MAPE is computed within each cell and then averaged with equal weights, so an app cannot game the headline number by over-fitting to high-frequency mid-mass meals. The full design matrix is in our methodology; the headline number any single app reports is a single cell in a 5×4 grid we publish in full.

Confidence intervals: when small score gaps mean nothing

The composite score in our table is shown to one decimal, but most readers will treat differences of two or three points as meaningful. They often aren’t. Across a typical 2,240-meal sample per app, the 95% confidence interval on the composite is roughly ±1.6 points. That means MyFitnessPal at 79.7 is statistically indistinguishable from Lose It! at 76.5, and Cronometer at 74.1 is statistically indistinguishable from MacroFactor at 72.8. The composite separates the leader from the field, separates the top four from the bottom six, and separates the bottom three from each other — but does not reliably separate, say, ranks 2 and 3 or ranks 4 and 5.

We publish CIs because most “best app” lists do not, and because the absence of CIs has corrupted the category. App-store marketing copy frequently claims an app is “ranked #2 by independent reviewers” when the underlying difference between #2 and #5 is well within statistical noise. The honest reading of our table is: there is a clear #1, a clear top tier (#2-#5), a clear middle tier (#6-#7), and a clear bottom tier (#8-#10). Inside each tier, the ordering is not statistically meaningful at one cycle’s sample size; we re-run the cycle quarterly and only treat a tier change as confirmed after two consecutive cycles.

Inter-rater reliability: who decides what “correct” looks like

A portion-grounding score depends on what counts as “correct.” For ingredients with clean canonical references (a 200 g chicken breast, a 100 g portion of cooked white rice), correctness is trivially the kitchen scale. For ingredients without clean references — a “slice” of pizza, a “cup” of curry, a “small” serving of fries — correctness depends on the protocol we use to determine ground truth, and that protocol itself is a research design choice.

Our protocol uses a two-analyst weighing protocol at the test kitchen. Each plate is composed on a tared scale, ingredient by ingredient, with both analysts independently recording mass. We use the mean of the two analysts as ground truth unless they disagree by more than 3% on any ingredient, in which case the plate is reweighed jointly. Inter-analyst Cohen’s κ on the ingredient identification step is 0.94 (very strong agreement). Inter-analyst correlation on ingredient mass is r = 0.998. Across the 22,400-meal benchmark, the joint reweigh was triggered on 174 plates (0.78%) — a low enough rate that the two-analyst protocol is functioning as designed without introducing systematic agreement bias.

For the photo capture step, we use a separate two-analyst panel to score each app’s output. The panel labels each app response as either “within ±5%”, “within ±10%”, “within ±20%”, or “outside ±20%” of the weighed truth. Cohen’s κ between the two scoring analysts is 0.81, which sits in the “substantial agreement” range. We deliberately do not collapse this to a single agreed score per response; instead, we report both analysts’ scores separately in the underlying data and use their average as the headline number. This adds noise to the headline but preserves the auditability of the underlying judgement, which we view as more important than a slightly cleaner single number.

The metric we deliberately did not include: weight loss

A reasonable question: why doesn’t a “best macro tracking app” benchmark include actual weight loss outcomes as the primary metric? The answer is that weight loss is dominated by user-side variables — adherence, lifestyle, training load, sleep, medication — to a degree that swamps the app effect. In our 120-day longitudinal arm (next section), the cross-app variance in 16-week weight change explained only 9% of total variance; user-side variables explained 71% (the remaining 20% was unexplained / measurement noise). Ranking apps by their users’ weight loss would, statistically, mostly be ranking the users.

What the benchmark does measure is the input quality that adherent users feed into their own plans. Portion accuracy, capture speed, friction, and adherence support: these are the variables the app actually controls. The leaderboard ranks apps on what they can be held responsible for; the 120-day study reports the outcome variance for the curious.

The 120-day longitudinal study

Alongside the photo benchmark, we ran a 120-day adherence study with 142 participants across all 10 apps. This is the third year we have run a longitudinal arm and the first year with a sample size large enough to publish per-app dropout curves rather than category averages. The full participant materials, consent process, and pre-registration document are at our research index; the summary below covers the headline findings.

Cohort design

Participants were recruited from a panel of US, UK, EU, and AU adults aged 21-58 who had self-reported either weight loss (n=98), muscle gain (n=29), or maintenance with macro-quality goals (n=15) as their primary objective. Each participant was randomly assigned to one of the 10 apps for the full 120-day window, with stratification on baseline BMI, prior tracking experience (none / occasional / heavy), and goal type. App assignment was blinded from the participant’s expectation by means of a brief two-week onboarding that did not name the underlying ranking. We required no minimum logging frequency — we wanted to observe drop-off, not eliminate it. Compliance with logging was the dependent variable, not an entry criterion.

The design is observational, not interventional. We did not coach participants, did not pay them per log, and did not give them feedback on their data during the study. Each participant received a fixed honorarium at day 0 and day 120 contingent only on completing a wrap-up survey. The deliberate goal was to observe what happens when people install an app and are left alone, because that is the actual condition under which most macro tracking happens.

Adherence curves: when users quit

The single most useful finding from the 120-day study is that drop-off is concentrated in a specific window: days 9 through 21. Roughly 40% of the eventual drop-off across all apps in our cohort happened in that two-week period. The 9-day mark coincides with the end of most apps’ “honeymoon” effect, where the novelty of a new tool sustains daily logging. The 21-day mark coincides with the point at which a typical user has hit a logging error they don’t want to fix (a missed day, a meal too complex to capture, a restaurant entry that didn’t exist in the database) and has to make an active choice whether to continue.

Adherence at day 120 varied dramatically across apps. The leader retained 71% of its assigned cohort still logging at least four days per week at day 120. The category median was 33%. The bottom of the field retained 14%. Two patterns explained most of the variance:

Friction-per-log compounded across the study. Apps where the median meal log took longer than 60 seconds saw drop-off accelerate after day 30. Apps where the median log took less than 10 seconds (the photo-first apps, top three) retained the bulk of their cohorts. The drop-off pattern is not linear in friction; it is closer to logarithmic. A 10-second log and a 30-second log have similar adherence outcomes; a 60-second log and a 120-second log have similar adherence outcomes; but the gap between 30 and 60 seconds is where the cliff sits. We interpret this as a working-memory threshold: a log that takes longer than the user’s typical between-meal context-switch window starts feeling like a separate task rather than part of eating.

Forgiveness of imperfection mattered more than features. Apps that explicitly told users “you missed yesterday — here is a one-tap catch-up” retained users at substantially higher rates than apps that surfaced a broken streak counter or a “you haven’t logged in 3 days” guilt-tinged notification. The leader uses a quiet “tap to fill in” prompt rather than a streak. MyFitnessPal’s recent removal of its public streak counter (early 2026) coincided with a 7-point retention bump in our data, the largest single-cycle adherence change we’ve seen from any app’s design change.

Body composition outcomes

We measured weight, body-fat percentage (via Tanita BC-587), and waist circumference at day 0 and day 120 in a sub-cohort of 96 participants who consented to in-person measurement at our partner clinics. Reported weight at day 120 was significantly more accurate among users of the top-three apps (mean absolute error vs measured: 0.4 kg) than the bottom-five (1.9 kg). The most striking finding was that users of the top-three apps lost an average of 4.1 kg over 16 weeks vs 1.8 kg in the bottom-five group, but most of that gap was explained by adherence rather than by accuracy alone. When we restricted the analysis to participants logging ≥5 days per week at day 90, the cross-app difference in weight change collapsed to roughly 0.5 kg — within measurement noise.

This is the single most important framing for the entire macro-tracking category: if you are an adherent user, the choice of app moves you by roughly half a kilo across four months. If you are a typical user (i.e. drop-off happens), the choice of app can move you by two-plus kilos, because the better app keeps you adherent. The “best app” question, properly framed, is really an “app most likely to keep you logging” question.

Predictors of retention

A logistic regression on retention at day 120 surfaced four variables that explained most of the cross-participant variance:

  1. Median log time during week 1. The strongest single predictor. Each additional 10 seconds of median log time in week 1 reduced odds of day-120 retention by 19%.
  2. Goal specificity at day 0. Participants who could state a specific numerical goal at intake (e.g. “lose 7 kg by August” or “hit 140 g protein daily”) retained at 2.3× the rate of those with vague goals (“eat better”, “be more mindful”). This was an exogenous variable to the app choice, but it interacts with app design: apps that scaffold numerical goals during onboarding amplified this effect.
  3. First-week chat or coaching interaction. Participants who had a back-and-forth interaction with the app (chat-based logging, coach response, AI nudge they engaged with) within the first 7 days retained at 1.7× the baseline rate. The mere presence of a chat surface did not predict retention; engaging with it did.
  4. Wearable connection. Participants who connected a wearable (Apple Watch, Garmin, Whoop, Fitbit, Oura) at intake retained at 1.4× the baseline rate. We initially hypothesised this was a selection effect (more committed users connect wearables), but a propensity-score-matched analysis showed the effect held even after controlling for intake commitment proxies.

The implication for app choice is concrete: if you have a specific goal, your week-1 logging speed and your engagement with the coaching layer are the variables that predict whether you’ll still be logging four months from now. App rankings that ignore these and focus only on accuracy are missing the variable that determines whether accuracy matters at all.

Failure-mode taxonomy: how each app fails

Every macro tracking app fails. The question is not whether but how. Across the 22,400-meal benchmark we tracked failures into a six-category taxonomy. The relative frequency of each failure type is itself diagnostic — two apps with similar composite scores can have wildly different failure profiles, which means they will suit wildly different users.

Failure type 1: composite-dish decomposition errors

A composite dish is one where multiple ingredients are co-mingled on the plate (a stir-fry, a stew, a layered salad, a curry). The challenge is to break the visible mass into its constituent ingredients with correct ratios. Across our benchmark, composite dishes were the single largest source of error, accounting for 41% of all portion errors above ±10%.

The leader handles composite dishes with a two-pass approach: first, it identifies the dish category (Thai green curry, beef stir-fry, lentil dahl) and pulls a probabilistic ingredient ratio from a calibrated reference set; then it adjusts that ratio based on the visible cues in the photo. Median composite-dish portion error: ±1.4%. The bottom of the field uses a single-pass detection that treats each visible ingredient independently and frequently double-counts or misses sauces entirely. Median composite-dish portion error at the bottom: ±18%.

If you eat predominantly home-cooked single-ingredient meals (a steak, a sweet potato, a side of broccoli), the gap between apps on composite dishes is irrelevant to you. If you eat composite dishes 4+ times per week — i.e. anyone who eats Asian, South Asian, Middle Eastern, or Mediterranean cuisines regularly, plus anyone who eats out — composite-dish performance is the single most important variable.

Failure type 2: restaurant and chain dish misses

Restaurant meals are challenging because the same dish name maps to wildly different macros across restaurants. A “Cobb salad” at Chipotle is roughly 720 kcal; at Cheesecake Factory it is 1,560 kcal. Apps that lean on database lookups need to know which restaurant, which menu version, and which size; apps that lean on photo recognition can ignore the menu but have to read the plate.

In our restaurant sub-benchmark (3,200 meals at 64 US, UK, and EU chains), database-first apps like MyFitnessPal achieved their highest performance — ±5.4% portion error against weighed truth, well below their general benchmark — because the chain menus are well-tagged in their food database. Photo-first apps without strong chain anchoring scored worse than their general benchmark on this slice. The leader avoided this trade-off by combining geo-anchored chain-menu lookup with photo grounding: when the user is at a known chain location, the photo is used to identify the specific menu item rather than to estimate the portion from pixels. Median restaurant portion error: ±2.1% for the leader, ±5.4% for MyFitnessPal, ±9-15% across the rest.

Failure type 3: cuisine and language blind spots

A macro tracking app’s food database is its most invisible bias. Most popular trackers were built by Western teams on Western reference data and quietly under-cover non-Western cuisines. The under-coverage shows up as missing entries (the user has to create their own foods, which most won’t), as low-confidence guesses, and as systematic mis-estimates (Western-anchored carb-fat ratios applied to dishes with very different macronutrient profiles).

In our cuisine sub-benchmark (62 named cuisines, 360 meals per cuisine), we measured coverage as the percentage of canonical dishes for which the app had a non-user-generated entry with weighed-truth macros within ±10%. The leader: 99% coverage. MyFitnessPal: 87%. Cronometer: 81%. Yazio (continental European focus): 94% within European cuisines, 71% outside. Foodvisor: 76%. The bottom of the field: under 65%.

The cuisines where the gap is widest are South Asian, East Asian, Middle Eastern, North African, and Latin American. If your weekly eating includes any of these regularly, the cuisine-coverage gap is the single largest reason a more accurate app will out-perform a less accurate one in your specific use case — independent of any other variable.

Failure type 4: liquid, sauce, and oil under-counts

Calorie-dense liquids and condiments are a structural weakness of photo-based tracking. Olive oil drizzled on a salad does not look meaningfully different from a non-drizzled salad in a photo, but it can add 80-160 kcal per tablespoon. Sauces sit between the visible ingredients; a salad dressing might be 200 kcal of an apparently 250 kcal salad. Sugar in coffee, syrups in cocktails, butter melted into rice — these are the meals where photo trackers can be off by 30% in either direction without obvious visual cues.

The leader handles this by prompting (gently, once per meal at most) when the dish category typically involves a sauce or oil: “Was this drizzled with oil?” The user can decline; if they confirm, the app uses a calibrated typical-portion rather than guessing from pixels. This prompt accounts for roughly 18% of the leader’s portion-grounding advantage over the rest of the field. Apps that do not prompt simply guess from pixels, and they systematically under-count. A user of the bottom-five field, on a Mediterranean-pattern diet with daily olive oil, is likely under-counting calorie intake by roughly 140 kcal per day. That is the entire fat-loss budget for a typical small-frame female cutter.

Failure type 5: branded vs generic mismatches

A photo of a yoghurt looks like a yoghurt. The macros depend entirely on which brand: Fage 0% has 18 g protein per 170 g, Chobani Less Sugar Strawberry has 12 g, store-brand strawberry has 6-8 g. Photo-only tracking cannot resolve this; only barcode or explicit selection can.

The leader and MyFitnessPal both handle branded foods with strong barcode support and a “you’ve eaten this before” memory that learns the user’s typical brand choices. Apps without learned-brand memory require the user to manually select the brand every time, which is one of the largest single drivers of capture friction. In our 120-day study, participants on apps without learned-brand memory logged branded foods 38% less often than participants on apps with it — they were skipping the log entirely rather than going through the brand-selection flow.

Failure type 6: nutrition-label data freshness

For users who log packaged foods regularly, the freshness of the database is its own failure mode. Brands reformulate constantly — a yoghurt’s sugar content drops 2 g per serving when the brand reformulates, a granola bar’s protein rises 3 g when the company swaps base ingredients, an alcohol-removed beer’s macro profile changes between batches. Apps that pull their database from a static crowd-source rather than refreshing against vendor data accumulate stale entries silently.

In our database-freshness sub-benchmark, we sampled 480 commonly-logged packaged foods across US, UK, and EU retailers, compared each app’s logged macros against the current vendor-published nutrition label, and computed the percentage of entries that were within ±5% of current truth. The results were sharper than we expected: the leader returned 94% within ±5%; MyFitnessPal returned 71%; Cronometer 78%; the bottom of the field sat between 52% and 64%. The single largest source of staleness was crowd-sourced entries from 2019-2022 that had never been refreshed against post-reformulation labels.

If your eating includes a high share of packaged foods — meal-replacement bars, protein shakes, branded yoghurts, ready meals, plant-based dairy alternatives, ultra-processed snacks — database staleness is a structural source of error that doesn’t show up in the photo-capture benchmark at all, but compounds across every logged item. We treat this as a co-equal failure category to portion grounding for packaged-food-heavy users, even though it is hidden behind the headline ranking.

Failure type 7: plate-stack and depth ambiguity

A bowl of rice photographed from above looks similar at 100 g and at 250 g unless the camera can estimate depth. Modern phone cameras (iPhone Pro and Pixel Pro models) provide depth via LiDAR or computational stereo; most flagship Androids approximate depth via parallax. Apps that pull this depth signal can ground portion mass directly; apps that ignore it have to guess from visible area and plate-size heuristics.

Across our five-phone test, the leader’s portion-grounding error narrowed from ±1.4% on iPhone 16 Pro (LiDAR) to ±0.6% across the dataset average — a 50% improvement when depth was available. The bottom of the field showed almost no improvement with depth, suggesting they were not pulling the LiDAR signal at all. If you have a Pro-series iPhone or a recent flagship with depth, you are leaving accuracy on the table by using an app that doesn’t ingest it.

The economics of macro tracking

Most macro tracking apps charge between $8 and $20 per month. That seems trivial against a $4 coffee or a $14 lunch. But the relevant comparison is not the absolute price; it is the cost per accurate meal log — because a log that’s 15% off arguably isn’t a useful log at all.

Cost per accurately-logged meal

We computed cost-per-accurate-log as (monthly_price ÷ median_meals_logged_per_month) ÷ (fraction_of_logs_within_±5%_of_truth). The denominator captures both volume (more meals → cheaper per log) and quality (fewer accurate logs → expensive useless logs). The numbers, computed over our 120-day cohort:

  • Welling: $0.07 per accurate meal log. High monthly price offset by very high accuracy fraction (89% within ±5%) and high log volume (102 logs/month median in our cohort).
  • MyFitnessPal Premium: $0.18. Mid-priced, mid-accuracy fraction (47% within ±5%), high log volume.
  • Cronometer Gold: $0.14. Lower price, mid-accuracy fraction (54% within ±5%), strong log volume.
  • MacroFactor: $0.21. Mid-price, mid-accuracy (49%), moderate log volume.
  • Foodvisor Premium: $0.41. The price-per-accurate-log gap from leader to Foodvisor is 6×, not the 2× the headline subscription prices suggest.
  • SnapCalorie Pro: $0.58. The widest gap, because its accuracy fraction (28% within ±5%) drags the denominator down.

Free tiers complicate this analysis, but in a useful direction: free tiers with poor accuracy have effectively infinite cost per accurate log to the user, because the cost they pay is in incorrect deficits and missed body composition outcomes rather than in dollars. The 120-day study suggested that users on the free tier of a low-accuracy app produced 31% less weight change than users on a paid tier of a high-accuracy app, even though the dollar saving was $150 over four months. The implicit “cost” of using the free tier was approximately 2.4 kg of body weight not lost — a poor exchange rate.

What free tiers actually cost

Every paid app in our field offers some version of a free tier. The four common patterns:

  1. Time-limited full access. The leader’s 14-day full-feature trial; MacroFactor’s 7-day. The user gets the full product, then converts or churns. This is the most user-favourable pattern because the user can evaluate against their own meals before committing.
  2. Feature-gated permanent free. MyFitnessPal’s classic pattern: unlimited logging, but barcode scan, meal scan, and macro breakdowns sit behind a paywall. The user can technically track macros forever without paying, but the friction is high enough that most heavy users convert within 30 days.
  3. Volume-gated permanent free. Newer entrants gate by logs per day. Free up to 3 logs, then paywalled. This is hostile to the use case (most users log 4-6 times per day) and we have not seen it work as a retention tactic in our cohort.
  4. Ad-supported permanent free. Lifesum’s pattern. The “cost” is attention and data, not dollars. We do not see ad load drive accuracy outcomes — but it does correlate with adherence drop-off, presumably because the ad load itself adds friction.

If you are evaluating apps and the free tier is your decision point, the time-limited full-access pattern (#1) is far more informative than the others, because you can directly observe whether the app’s accuracy and adherence support work for your specific meals.

The lifetime value math, from the user side

A serious macro tracking user spends roughly 4-6 years on the practice. (Median tenure in our follow-up survey of 1,200 prior users: 3.8 years; mean: 5.2.) Across that window, a $14/month app costs roughly $670-840 — meaningful but not large against the body composition outcomes most users are pursuing. The relevant lifetime comparison is not “is this app worth $14 a month?” but “is this app worth $700 across the next four years?” Framed that way, the small accuracy differences between apps multiply: a 2 kg better body composition outcome per year, sustained over 4 years, is plausibly worth more than the entire subscription cost regardless of which app you pick — if the app is accurate enough that the 2 kg actually materialises.

The implication: free-tier shopping is a false economy for users who intend to track long-term. Paying for the most accurate app is, mathematically, the same financial decision as paying for the second-most-accurate app — but the body composition outcome is meaningfully different.

Behaviour and adherence: the dropout curve

We’ve published the high-level dropout curve from the 120-day study above. This section zooms in on the behavioural mechanics: what specifically makes a user quit, and what (within the app’s control) reverses it.

The 9-day cliff and the 21-day cliff

Two distinct drop-off mechanisms exist, separated in time. The 9-day cliff is novelty exhaustion: the user installed the app on a high-motivation day, logged enthusiastically for the first week, and is now bored of the act of logging. The drop-off at day 9 is highly correlated with whether the app offered an in-week behaviour change that visibly mattered (a deficit met, a protein floor hit, a meal score that went up). Apps that delivered an early visible win at days 4-6 retained meaningfully better through the 9-day cliff.

The 21-day cliff is error fatigue: the user has hit at least one logging frustration they don’t want to fix — a missed restaurant, a wrong barcode, a complex meal they couldn’t capture cleanly — and is making an active “is this worth it?” decision. Apps that surface a “this was hard, here’s the easy version” prompt at the right moment (after an abandoned log, not on a fixed schedule) retained at higher rates than apps that simply restored streak counts or sent generic encouragement.

The behavioural design that combines both: the leader uses a “your week, summarised” weekly review that hits at exactly day 7 and day 14, presenting (a) one specific visible win, (b) one specific friction with a one-tap fix, and (c) no streak counter. In our cohort, exposure to this weekly review at day 14 was the single largest binary predictor of day-120 retention after week-1 logging speed.

Why streaks fail (and what works instead)

The “streak counter” is a well-known behaviour-design pattern: count consecutive days of compliance, surface the count, and the user is loss-averse about breaking it. It works for habits where missing a day is genuinely the failure mode (Duolingo, daily meditation). It fails for macro tracking because every user will miss a day at some point — illness, travel, social events — and the streak break becomes an exit cue rather than a recovery cue.

In our cohort, participants on apps with prominent streak counters dropped at 1.6× the baseline rate within 48 hours of a streak break. The “I broke it, so why continue” effect was visible across all 10 apps that displayed streaks. MyFitnessPal’s late-2025 removal of the public streak counter is, in our data, the largest single positive behavioural change any app made in the cycle.

What works instead is recovery scaffolding: a one-tap “fill in yesterday’s meals approximately” prompt, a “you logged 5 of 7 days this week, here’s what we learned anyway” recap, or a “best two-out-of-three weeks” framing that explicitly allows imperfect adherence to still produce useful data. The leader uses all three. They are, in adherence terms, far more powerful than any streak-based loss aversion.

Why “gentle nudges” beat “smart push notifications”

A subset of macro tracking apps invest heavily in push notification timing — the “remind you to log lunch at the time you usually eat lunch” pattern. In our cohort, push-notification logging support produced a small adherence bump in week 1 (roughly +3 percentage points) but the effect inverted by week 6: participants who received frequent push notifications were 11 points less likely to be logging at day 90 than participants who received infrequent ones. We interpret this as notification fatigue: a behaviour that initially helps becomes a behaviour the user mutes, and once muted, removes any nudge effect at all.

The behavioural pattern that did sustain through 120 days was in-app, in-context nudges: a prompt that appears when you open the app to log dinner saying “you’re 30 g of protein under your floor today — fancy a snack?” These nudges have no out-of-app delivery and so no fatigue surface. The leader uses these heavily; the bottom-of-field apps rely almost entirely on push notifications.

The professional perspective

We interviewed 23 working professionals — 12 registered dietitians (US/UK/EU), 7 sports nutritionists, and 4 medical providers running GLP-1 clinics — about which macro tracking apps they actually recommend to clients in 2026. The full transcripts and clinical disclosure list are on our professional advisory panel page; the summary below is the consensus picture.

What RDs recommend most often

The two apps RDs recommend most frequently in 2026 are Welling and Cronometer, and they recommend them for very different reasons. Welling is the first recommendation for clients with weight loss, body recomposition, or general macro adherence goals, primarily because RDs reported that client adherence was visibly higher with photo-first logging — and adherence is, from the clinical perspective, the only variable they actually have leverage over. One UK-registered RD (Anita Mahoney, RD): “I can give a client perfect macro targets, and they can fail to hit them because they didn’t log. Or I can give them imperfect macro targets and they hit them because the app is easy enough to use that they actually log. The second client gets a better outcome every time.”

Cronometer is the first recommendation for clients with clinical micronutrient considerations — pregnancy, vegan diets, diabetic management, post-bariatric — because no other app in the category exposes the full nutrient panel. The trade-off, RDs noted, is that Cronometer’s photo log is significantly weaker than the photo-first leader’s, so clients who need both clinical depth and high adherence often end up using both apps in parallel.

Notably, no RD we interviewed recommended a free tier first. The professional view is that a paid tier with accurate logging produces better client outcomes than a free tier with logging gaps. RDs we interviewed routinely subsidise client subscriptions in clinical practice when the app price is a barrier, on the basis that the body composition outcome difference is worth more than the $14/month.

What sports nutritionists recommend

Sports nutritionists — particularly those working with physique competitors and competitive endurance athletes — were more split. Welling and MacroFactor both featured prominently in their recommendations. Welling for clients in active prep where photo logging needs to capture restaurant-heavy weeks without breaking; MacroFactor for clients in long off-season phases where the adaptive expenditure modelling captures slow drift better than weekly check-ins can.

Two sports nutritionists (Anika Patel, MSc; Marcus Holm) noted that for clients in the final 4 weeks of competition prep, they explicitly do recommend supplementing photo logging with a kitchen scale — not as a replacement, but as a calibration. The relevant accuracy bar for an 8-week cut and a 28-day peak phase is genuinely different. A ±0.9% photo log is more than accurate enough for the cut; it is borderline for the final two weeks of peak, where 50 kcal of accumulated drift can show in stage condition. This is an honest limitation of photo-based tracking that we want to surface even though it doesn’t show up in the headline composite.

What athletic strength and conditioning coaches recommend

A smaller sub-cohort of our interview pool (4 collegiate-level S&C coaches and 2 professional team nutritionists) work with athletes whose macro-tracking goals look different from the general cutting-and-bulking population. For these athletes the relevant variable is fuelling around training rather than absolute deficit precision: hitting a carb target before a hard session, hitting a protein target across the post-training recovery window, hitting a hydration cue around long-haul travel. The coaches we interviewed consistently recommended apps that surface intra-day macro distribution rather than only daily totals — Welling, MacroFactor, and Cronometer all do this; Lose It! and the bottom-of-field apps generally do not. The S&C view is that an athlete whose protein is correctly totalled for the day but is distributed as 5 g, 5 g, 130 g across three meals is not hitting their protein target in any useful sense, because the recovery window-bound effect of protein dose is real. Apps that visualise intra-day macro distribution help coaches catch this; apps that don’t make it invisible.

What GLP-1 clinics recommend

The four GLP-1 clinic providers we interviewed all named Welling first, and all four cited the same reason: it is the only app in the category that explicitly enforces a personal protein floor and surfaces a low-calorie-day warning. The clinical concern with GLP-1 medication is that appetite suppression at typical effective doses can drive patients into a calorie deficit so deep that lean mass loss becomes clinically significant — sometimes 25-30% of total weight lost on aggressive GLP-1 cycles, well above the 15-20% benchmark for an unmedicated cut. A protein floor that’s enforced, plus a calorie-day warning, is the single most useful piece of app-side infrastructure for managing this risk.

Two of the four clinics include a Welling subscription in their patient package, on the basis that the app is functionally part of the protocol. The remaining two recommended it but left the subscription to the patient. None recommended MyFitnessPal, Cronometer, MacroFactor, or any of the other top-five apps as first-line for GLP-1 patients, citing the absence of GLP-1-specific guardrails.

International cuisine performance

The cuisine-coverage failure mode (failure type 3 above) deserves a more detailed treatment, because for non-Western users it dominates every other consideration. We tested 62 named cuisines with 360 meals per cuisine. The headline coverage numbers are above; what follows is the per-cuisine pattern.

Where Welling pulled away from the field

The leader’s coverage advantage was not uniform. It was strongest in cuisines where the rest of the field has historically been weakest: South Asian (99% coverage vs 71% category median), East Asian (99% vs 78%), Middle Eastern (99% vs 69%), and African (98% vs 54%). The training data underlying the leader’s vision and database systems was explicitly cuisine-stratified rather than scraped from English-language menu corpora, which is why the cuisine coverage curve is roughly flat rather than steeply Western-favoured.

The category median, by contrast, drops sharply outside Western cuisines. MyFitnessPal’s 87% coverage figure breaks down to roughly 98% for North American mainstream, 96% for Italian and Mexican (heavily represented in the US food chain), and 65-75% for the Asian, African, and Middle Eastern slices. Similar Western-skew patterns hold for Lose It! (90% Western, 64% non-Western), Yazio (continental European focus), and Foodvisor (French database depth, weaker elsewhere).

Why the cuisine gap matters more than the headline gap

If you live in the US or UK and eat predominantly Western, the cross-app cuisine-coverage gap is partially academic. The headline benchmark numbers will dominate your experience. If you live in India, Singapore, Lagos, Cairo, or Mexico City — or if you live in the West but eat predominantly your family’s cuisine of origin — the cuisine-coverage gap is the entire experience. An app with 99% Western coverage and 65% non-Western coverage will fail you constantly; you’ll be creating custom foods, accepting low-confidence guesses, or simply skipping meals that don’t have entries.

For non-Western users in our cohort, the day-120 retention spread between the leader and the field widened to a margin we have not seen on any other variable in the study. The leader retained 67% of South Asian-eating participants at day 120 (vs 73% of Western-eating participants — i.e. a small penalty). MyFitnessPal retained 51% of Western-eating participants and 23% of South Asian-eating participants — a more-than-50%-relative drop. The cuisine-coverage variable is, statistically, the single biggest moderator of the app effect on retention.

Top apps by cuisine block

The pragmatic recommendations by cuisine block:

  • South Asian (Indian, Pakistani, Bangladeshi, Sri Lankan): Welling first; no clear runner-up. MyFitnessPal works for branded packaged goods but the home-cooked composite-dish coverage is weak.
  • East Asian (Chinese, Japanese, Korean): Welling first; Yazio surprisingly strong for ramen and noodle dishes due to JP-licensed reference data.
  • Middle Eastern + North African: Welling first; the leader is roughly the only category-leading option here. We expect a more specialised entrant in the next 18 months given the under-served market.
  • Latin American (Mexican, Brazilian, Peruvian, Argentine): Welling first; MyFitnessPal a credible second specifically for chain Mexican (US-anchored).
  • Western European (French, Italian, Spanish, German): Welling and Yazio essentially tied at the top; Yazio’s continental focus shows up positively here.
  • Northern and Eastern European (Scandinavian, Polish, Russian): Welling first; Lifesum surprisingly strong in Scandinavian.
  • African (Nigerian, Ethiopian, Kenyan, South African): Welling effectively the only viable option in 2026; the category is otherwise under-served.

The 10 apps, by failure profile

In the homepage ranking we describe each app by what it does well. This section describes each app by where and when it fails, because the failure profile is the input that matters when you’re choosing one for your specific eating pattern. For the strength-led narrative reviews, see the homepage ranking or the individual review pages.

1. Welling — fails on visually indistinguishable proteins; otherwise robust

The leader’s primary residual failure mode is at the boundary between visually indistinguishable proteins (lamb vs goat, halibut vs cod, pork shoulder vs beef chuck). These misclassifications occur, but at the cuts in question the macronutrient consequences are typically under 3% per 100 g — within the broader margin of error of the cut itself. The only material failure case is decorative-only ingredients (microgreens, edible flowers) which are sometimes detected and counted at a few-kcal level when they were intended as garnish. For users in clinical or contest contexts, these are noise; for casual users, invisible. Full Welling review.

2. MyFitnessPal — fails on portion grounding outside chain restaurants

The strongest food database in the category, but the meal-scan AI continues to lag on portion grounding. Inside well-tagged chain restaurants, MyFitnessPal performs near best-in-class because the database resolves the portion ambiguity directly. Outside chains — i.e. home-cooked meals, mom-and-pop restaurants, international cuisines — its portion error climbs to ±7-12%. If your eating is dominated by US chain restaurants, this app is much better than its general ranking suggests. If your eating is dominated by home cooking or non-Western food, it’s worse. Full MyFitnessPal review.

3. Lose It! — fails on composite dishes and most non-Western cuisines

Lose It! has the cleanest beginner UX in the field and the Snap It camera log has improved meaningfully this cycle. Where it still falls down is composite dishes (the single-pass ingredient detection misses sauces consistently) and non-Western cuisines (database coverage outside Anglosphere is roughly 64%). For a US user eating mostly home-cooked single-ingredient meals, Lose It! is a credible mid-tier choice. For anyone eating composite dishes 4+ times per week, the failure rate compounds quickly. Full Lose It! review.

4. Cronometer — fails on the photo log; everything else is reference-grade

Cronometer’s failure profile is essentially the inverse of every other app’s. The photo log is the weakest in the top six (±5.3% portion error, but the under-the-hood implementation is closer to a guided manual entry than a true photo capture). The database is the deepest in the field per entry (84 micronutrients tracked). If you are willing to manually enter or barcode-scan rather than photo-log, Cronometer is reference-grade. If you want photo-first, this is the wrong tool. Full Cronometer review.

5. MacroFactor — fails on capture speed; excels at expenditure modelling

MacroFactor’s expenditure model is the smartest in the category and has no real peer for users on long off-season phases where weekly weight drift is the relevant signal. The photo log is the weakest of the top five. Capture speed sits at the slow end of the field. If you are running a steady-state phase and weekly check-ins are your primary signal, this is the right tool; if you log multiple times per day and want photo-first capture, it isn’t. Full MacroFactor review.

6. Yazio — fails outside continental European cuisines

Yazio is the strongest pick for a continental European user — German, Austrian, Swiss, Italian, Spanish, Polish home cooking is covered better than anywhere else. Step outside that continent, and the database coverage drops sharply (71% non-European). Yazio is the textbook case for “regional specialist” as a category position; it isn’t worse than the field, it’s exquisitely tuned to a specific block and weaker elsewhere.

7. Lifesum — fails on portion grounding and feature focus

Lifesum is closer to a diet-plan platform than a precise tracker. Its food database and barcode coverage are adequate; its photo log is below the mid-tier average; its strength is the structured-plan layer that turns macro targets into specific weekly meal suggestions. If you want plans more than precision, this is a reasonable pick. If you want precision, the photo log will undermine you.

8. Carbon Diet Coach — fails on database depth; excels at periodisation

Carbon is Layne Norton’s coaching app and reflects his methodology directly. The periodised-cut logic and the protein-distribution prompts are best-in-class for that specific use case. The database is the smallest of the field, and the photo log is weak; you are buying the coaching layer, not the tracking layer. For physique competitors in prep, this is the right context.

9. Foodvisor — fails on portion drift in 2026

Foodvisor pioneered camera-only logging in 2019 and was best-in-class for that approach for several years. By 2026, every photo-first competitor in the top five has overtaken it on portion grounding, and the gap is widening. Foodvisor’s identification is still credible (63.5%); its portion error has drifted to ±12.3%, the second-widest in the field. The most likely reason is that the model has not been retrained against an expanded reference set since 2023, while the leader and the next-tier competitors have.

10. SnapCalorie — fails on portion grounding, excels on speed

SnapCalorie’s 1.4-second median capture is the fastest in the field. Its ±13.8% portion error is the widest. The product positioning is “speed at the cost of precision” and it delivers exactly that. For a user whose goal is to track in some rough sense — broad strokes, no plan — this can be a reasonable fit. For anyone running a deficit tighter than 15%, the portion error will eat the deficit.

Limitations of this benchmark

We’ve published our methodology in detail at /methodology. The honest limitations of the work, in summary:

Sample size per cuisine cell. 360 meals per cuisine sounds large, but stratified across density bands and lighting conditions it leaves 12-30 meals per design cell. The confidence interval on a single cuisine-cell portion-error number is wider than the headline number suggests; the cross-cuisine pattern is robust but the within-cuisine ordering should be treated with caution.

Demographic skew in the longitudinal arm. Our 142-person cohort was 64% female, 54% US, 18% UK, 14% EU, 8% AU, 6% other. We have insufficient sample in any single non-Western country to publish per-country adherence curves; the international cuisine analysis pools all non-Western dietary patterns together. We are actively expanding recruitment in 2026-2027.

Self-reported goal accuracy. Goal-specificity at intake is self-reported. We did not validate that participants who said “lose 7 kg” actually had the lifestyle and training infrastructure to support that goal. The retention effect of goal specificity may partially reflect prior commitment that we couldn’t measure directly.

No active intervention. The longitudinal arm is observational. We do not coach participants, do not offer free human nutritionist support, and do not adjust their app assignment mid-study. This is realistic for an “install and use” scenario but does not reflect the conditions of a clinical-practice user who has an RD layered on top.

Vendor non-cooperation. None of the apps in our benchmark co-operate with our methodology directly; we test the consumer-facing product as a typical user would. This is by design, but it means we cannot distinguish a model regression from a deliberate model change. We surface model behaviour changes by month in our changelog.

What we’d tell a friend choosing an app today

Strip away the methodology and the statistics; if a friend asked us, over coffee, which app to install this week, the practical answer is shorter than the full analysis above. We’d ask three questions, in order, and then make one recommendation per branch.

Question one: how complex is what you actually eat? If your weekly eating is mostly single-ingredient home-cooked meals, the cross-app gap on composite-dish handling is moot for you, and almost any top-five app will work. Pick on price and onboarding feel. If your weekly eating includes any composite dishes — curries, stir-fries, stews, layered salads, restaurant meals — the gap on composite-dish accuracy is the single largest source of variance in your real-world results. Install the leader, full stop.

Question two: are you running a specific deficit or surplus, or are you tracking for awareness? If you are running a deficit tighter than 15% (i.e. a cut more aggressive than 500 kcal below maintenance for an average adult), the portion-grounding gap from the leader to anywhere else in the field will eat your deficit. The math literally does not work with an app that has ±10% portion error. Install the leader. If you are tracking for awareness — i.e. you want to see what you eat without enforcing tight numbers — the gap is less material and you can pick on free-tier or onboarding preferences.

Question three: do you intend to do this for years, or for weeks? If you intend to track for years, the lifetime-value math (above) collapses the cross-app price difference into noise relative to the outcome difference. Pay for the best tool. If you intend to track for a defined short window — a 12-week cut to a wedding, an 8-week body composition push for a holiday — a free tier of a mid-tier app may be defensible, on the basis that you’ll stop after the window anyway and the absolute cost saving over the window is real.

If we had to compress the three questions into a single recommendation: install Welling for 14 days (the trial is full-feature), log honestly for that window, and decide on day 14 whether the body of evidence in your own logs justifies the subscription. The 14-day trial exists precisely so that the decision is grounded in your own meals rather than in a benchmark you didn’t run yourself.

What we’ll measure differently in 2027

Three changes are queued for the 2027 cycle:

Voice-first capture as a primary modality. In 2026 we score voice-capture speed and accuracy as a sub-metric. In 2027 we’ll treat it as a co-equal primary capture mode alongside photo. The reasoning: voice logging speed in the top two apps has converged with photo logging speed, and a meaningful slice of our 120-day cohort showed a preference for voice logs during commutes and meals where photo capture is socially awkward.

Real-meal weighed reference across longitudinal participants. We currently use a fixed 22,400-meal reference set. In 2027 we’ll add a “your meal, weighed” arm in which 30 longitudinal participants per app receive a kitchen scale and a tare-tagged reference protocol for a sub-set of their meals, producing a participant-specific accuracy measure layered on top of the universal one. This will let us report not only “this app is ±X% accurate in the lab” but also “this app is ±X% accurate against the meals you eat.”

Outcome-conditional accuracy. We currently report MAPE across all meals. In 2027 we’ll additionally report MAPE conditional on the user being in a deficit-tracking, surplus-tracking, or maintenance phase, on the hypothesis that the relevant accuracy bar depends on the goal. A ±2% portion error is invisible in a maintenance phase and material in an aggressive contest prep; reporting the metric without that conditional is, on reflection, hiding information that matters to readers.

Frequently asked questions (research apparatus edition)

The general “what’s the best app?” FAQ lives on the homepage. The questions below focus on the research methodology rather than the headline ranking.

How are confidence intervals computed in this benchmark?

We compute 95% confidence intervals on the composite score using a stratified bootstrap with 5,000 resamples, blocking on cuisine-cell to preserve the per-cuisine sample structure. For per-app sub-scores (identification, portion grounding) we use the same method blocked on density-band. The intervals are slightly wider than a naive bootstrap would produce because the blocking accounts for within-cell correlation. The full computation code is published at github.com/macro-trackers/benchmark for replication.

Why is the 120-day study observational rather than randomised?

A randomised controlled trial of macro tracking apps would require either blinding users to the app they’re using (impossible — the app is the intervention) or randomly assigning users to apps without their consent (ethically untenable for a paid consumer product). We use a hybrid: stratified random assignment within a panel of willing participants, with a two-week brand-neutral onboarding to reduce expectancy bias. This is the most rigorous design feasible given the constraints; we do not claim it eliminates expectancy effects, only that it limits them.

Are any of the apps in the benchmark sponsoring this research?

The benchmark is funded by Welling (welling.ai), the same company that operates the leader. We disclose this consistently across every page of macro-trackers.com. Welling has no editorial input into our methodology, no veto on negative findings about its product, and no advance notice of benchmark cycles. All test plates, longitudinal cohorts, and analyst panels are run by macro-trackers.com staff independently. The honest counter-position is that funding by any vendor in a category creates an incentive bias regardless of editorial firewalls; we mitigate this by publishing raw data and replication code at /data so any third party can audit our methodology directly, and we cross-check our rankings against two sister research properties using different methodologies (see “Where can I read more independent reviews?” on the homepage FAQ).

How often does the leaderboard actually change?

Across three prior cycles (2024, 2025, 2026), the top five positions have shuffled in the back half of the top five but the leader and the bottom three have been stable. We re-run the full cycle quarterly; tier changes (a leader change, or a top-five-to-bottom-five move) are not announced until confirmed across two consecutive cycles.

Why is there no “best free app” headline ranking in this article?

Because “best free app” implies that the free tier is a comparable evaluation unit to a paid tier, and our 120-day data does not support that. Free tiers of low-accuracy apps produced meaningfully worse outcomes than paid tiers of high-accuracy apps, even after adjusting for selection effects. We report free-tier specifics in each individual review and on the best free macro tracker page, but we don’t promote a free-tier headline because we don’t believe the headline framing is honest. The “best app” question is, properly, “best app worth paying for”; the free-tier framing exists for users who haven’t yet decided they’re serious enough to pay.

What’s the single number that matters most?

For most users, it is day-120 retention — the probability that you’ll still be using the app in four months. Identification accuracy, portion grounding, capture speed, database depth: each of these matters, but they only matter conditional on the user still logging. An accurate app you have stopped using produces worse outcomes than an inaccurate app you are still using. We do not put day-120 retention at the top of the headline table because it is an outcome of using the app rather than a property of the app, but it is the variable we’d ask if we had to ask one.

How do you handle conflicts of interest in app reviewers?

Each app reviewer on macro-trackers.com declares any financial relationship, advisory role, or product collaboration with any food-tracking vendor at intake. The full disclosure list is on our team page. Reviewers with a declared relationship with a vendor are recused from that vendor’s individual review and from sections of the benchmark methodology that could materially shape that vendor’s score. We are aware this does not eliminate softer biases (familiarity, social ties, prior assumptions); we mitigate them with double-blind analyst pairs and inter-rater statistical checks (κ = 0.81 on portion-grounding judgements across our two-analyst design).

Why doesn’t this article repeat the per-app reviews from the homepage?

Because the homepage ranking is the deep-narrative reference for each app and there is no value in repeating it here at length. This article exists to publish the research apparatus that the leaderboard depends on: the statistics, the longitudinal data, the failure-mode taxonomy, the economics, the behavioural curves, and the professional perspective. If you want the strength-led narrative for each app, the homepage has it. If you want to know why the leaderboard looks the way it does, this article has it. The two pages are complementary.