ingredient investigation
Yuka vs. Bobby Approved vs. Think Dirty: What Each App Actually Checks
Yuka, Bobby Approved (EWG), and Think Dirty all score your cleaning products — but they check different things, using different data. Here's the real difference.
You’re standing in the cleaning aisle, phone out, scanning a barcode with one app while your other hand hovers over a second bottle because the first app just gave you a yellow “moderate” and you don’t trust it. So you open a second app. It says something different. Now you’ve got two scores, no consensus, and a cart half-full of things you’re less sure about than when you walked in.
If that’s happened to you, you’re not doing it wrong. Yuka, Think Dirty, and Bobby Approved (built on EWG’s data) are not measuring the same thing, using the same data, or applying the same math — so getting different answers from different apps isn’t a glitch. It’s three different methodologies looking at the same ingredient label and reaching independently reasoned conclusions.
Let’s actually open the hood on each one, because knowing what each app is doing under the surface is worth more than any single score it hands you.
The short answer
None of these three apps tests your physical bottle in a lab. All three read the printed ingredient list and cross-reference it against public hazard databases — they differ in which databases, how many, and what formula they use to turn “these ingredients exist” into a single number or letter. Yuka scores off your single worst ingredient. Think Dirty runs a 0–10 “Dirty Meter” that auto-penalizes undisclosed fragrance. EWG (the data behind Bobby Approved and its own Guide to Healthy Cleaning site) grades A–F by pooling 15 separate databases, and treats “no data available” as a caution flag, not a clean bill of health.
What Yuka is actually doing when it scans your bottle
Yuka is the one most people recognize by its bright green, yellow, or red dot. According to Yuka’s own help center, a household or cosmetic product’s score — on a 0 to 100 scale — is calculated from the highest-risk ingredient present, not an average across the whole formula. If your all-purpose spray has nineteen mild ingredients and one flagged as high-risk, Yuka’s score reflects that one ingredient, capping the whole product below 25 out of 100.
That’s a meaningful design choice worth understanding, not just accepting. It means Yuka is built to catch a single bad actor hiding in an otherwise clean formula — which is useful, because that’s exactly how a lot of “natural-sounding” products are formulated. But it also means two products can share the same low score for very different reasons: one because nearly everything in it is a concern, another because of a single flagged preservative in an otherwise unremarkable mix. The number alone doesn’t tell you which situation you’re looking at — you have to tap through to the ingredient itself.
Yuka’s scientific backing draws on a specific, named set of regulatory and scientific bodies: the SCCS (EU’s Scientific Committee on Consumer Safety), ECHA (European Chemicals Agency), the US EPA, AICIS, ANSES (France’s food, environmental, and occupational health safety agency), and IARC — the same body that classifies substances as carcinogenic. It leans European by design, since the app originated in France, which is worth knowing if you’re comparing its verdicts against a US-built tool.
Think Dirty’s “Dirty Meter” — and why fragrance gets an automatic penalty
Think Dirty runs a different scale entirely: 0 to 10, branded the Dirty Meter. Per Think Dirty’s published methodology, every ingredient is separately evaluated across three specific health concerns — carcinogenicity, developmental and reproductive toxicity, and allergenicity/immunotoxicity — with each rated 4 to 10 depending on how strong the supporting evidence is (moderate, strong, or conclusive).
The detail that surprises people most: Think Dirty automatically scores any product listing undisclosed “fragrance” a 7 or higher, on the strength of the documented allergen and toxin risk baked into that legally vague catch-all term alone. That’s a direct, deliberate response to the same labeling loophole I’ve written about before — under FDA rules, a manufacturer can list “fragrance” without naming a single compound inside it, and Think Dirty’s team decided that opacity itself is worth penalizing, independent of what’s actually in the blend.
Think Dirty’s data sources skew broader-continent than Yuka’s: Health Canada, Environment Canada, and Environmental Defence Canada on the Canadian side; FDA, the National Toxicology Program, CDC, EPA, the California Safe Cosmetics Program, and the Cosmetic Ingredient Review (CIR) on the US side; and the EU Health Commission and Denmark’s Ministry of Environment on the European side. Its Chemistry Team and Advisory Board includes people with backgrounds at Health Canada, FDA, and EPA — which is a real, verifiable credential, not a marketing claim.
Bobby Approved and EWG’s Guide to Healthy Cleaning — the deepest data pool of the three
“Bobby Approved” is EWG’s consumer-facing brand for its ingredient-safety work, built on the same underlying research as EWG’s Guide to Healthy Cleaning. It grades products A through F, and per EWG’s own published methodology, that grade combines two separate things: how hazardous the evidence says an ingredient is, and how much evidence actually exists.
That second part is the detail that sets EWG apart from the other two: an ingredient with no available hazard data defaults to a C, not an A. EWG is explicit that “no data” is not the same claim as “proven safe,” and it refuses to reward the absence of research with a clean grade. Evidence that an ingredient doesn’t pose a hazard moves it up to a B or A. Evidence that it does moves it down to a D or F. A middling C simply means: nobody’s proven it’s dangerous, and nobody’s proven it’s fine either.
EWG’s grading pulls from an unusually wide net — 15 separate government, academic, and industry hazard and toxicity databases, including the EU’s Globally Harmonized System hazard classifications, the Association of Occupational and Environmental Clinics’ list of asthma-causing substances, and California’s Proposition 65 list of chemicals the state has determined cause cancer or reproductive harm. As of its published methodology, EWG’s Guide to Healthy Cleaning has graded more than 2,000 household cleaning products this way.
What that actually looks like, side by side
More sources isn’t automatically “more correct” — it’s a proxy for how wide a net each app casts, not a scoreboard for who’s right. A narrower, more targeted set of authoritative regulatory bodies (Yuka’s approach) can still catch the ingredients that matter most. But it does explain a lot of the disagreements you’ll see between apps on the same bottle: they’re not just weighting the same facts differently, they’re often drawing from genuinely different pools of facts.
Where they agree, and where they don’t
In my own side-by-side scans, the three apps mostly agree on the obvious offenders — undisclosed “fragrance,” certain formaldehyde-releasing preservatives, ingredients with a Prop 65 listing. That’s not a coincidence: those are the ingredients with the strongest, most conclusive evidence behind them, so every methodology converges on the same verdict.
Where they diverge is the low-dose, thinly studied middle ground — an ingredient with one animal study and no human data, or a preservative used at a concentration below where most of the concerning research was conducted. That’s precisely the zone where the underlying science itself is thinnest, so it makes sense the apps built on top of that science would disagree too. A different score there isn’t one app being “wrong” — it’s three different teams making three different judgment calls about how to treat genuine scientific uncertainty.
The micro-lesson: an app is a research shortcut, not a verdict
Here’s the one thing worth remembering the next time your phone gives you a color, a number out of 10, or a letter grade: every one of these tools is doing exactly what you could do yourself with the ingredient list and a few hours of database searching — they’re just doing it in two seconds. That’s genuinely valuable. It is not the same thing as a lab certifying your specific bottle is safe.
The most useful habit isn’t picking one app and trusting it blindly. It’s noticing which specific ingredient triggered a low score, and checking whether that flag shows up independently in a second app. Two different methodologies landing on the same ingredient, using different data pools, is a far stronger signal than any single score standing alone.
The simpler path
There’s an easier way to sidestep the whole scan-and-compare ritual: buy something with so few ingredients, and so little to hide, that there’s nothing left for an app to disagree about.
Ecolosophy’s All-Purpose Cleaning Concentrate is built plant-based, with no undisclosed “fragrance” trade-secret blend for Think Dirty to auto-flag and no mystery preservative for Yuka to score against its worst-ingredient rule. It’s small-batch, made with care, and one bottle makes 100+ spray bottles — so instead of scanning a new bottle every few weeks, you scan it once. Our co-founder Elizabeth is a PhD scientist and mom who reads these same databases for a living; the formula was built to hold up under exactly this kind of scrutiny, not to dodge it.
Frequently Asked Questions
Which app is most accurate: Yuka, Think Dirty, or Bobby Approved? None of them is running a lab test on your bottle — all three read the ingredient label and score it against public hazard databases. EWG’s Guide to Healthy Cleaning has the broadest data pool (15 pooled databases) and the most conservative default for unrated ingredients. Yuka is transparent that it scores off the single highest-risk ingredient. Think Dirty is most aggressive about penalizing undisclosed fragrance. Cross-checking two of the three on a borderline product beats trusting any single score.
Why do the apps sometimes give the same product different scores? They use different databases and different scoring math. Yuka leans on European bodies and scores off the worst single ingredient. Think Dirty pulls from about nine North American and European sources and auto-penalizes undisclosed fragrance. EWG pools 15 databases and grades on both hazard evidence and how much data exists at all.
Does an app give a bad score just because an ingredient hasn’t been studied? For EWG, yes by design — an unrated ingredient defaults to a C, not an A, because “no data” isn’t the same claim as “proven safe.” Yuka and Think Dirty handle unrated ingredients differently, but every app is working with real, permanent gaps in public toxicity data.
Do these apps actually test the product, or just read the label? They read the label. Each ingredient name is matched against a pre-built hazard database, and the score is calculated from that match. The score is only as good as what’s printed on the bottle — and a manufacturer can legally list a mystery blend as simply “fragrance” without naming a single compound.
What should I actually do with a low score? Treat it as a lead, not a verdict. Look at which ingredient triggered it, then check if it’s a documented concern (EPA, IARC, or Prop 65-listed) or a low-confidence “no data” flag. Two apps flagging the same ingredient independently is a stronger signal than one.
For more on the labeling loophole that Think Dirty automatically penalizes, read what “fragrance” actually hides. For the disinfectant ingredient these apps flag most often after fragrance, see quats in cleaning products, explained. And for a faster way to read any label yourself, without opening an app at all, see reading an ingredient label in 60 seconds. Browse the full Is It Toxic? series for individual product breakdowns, or shop the collection of concentrates built to score clean on all three apps at once.
#cleanwithlove #ecolosophy #nontoxichome #detoxyourlife #plantbasedliving
Sources cited
- Yuka Help Center — How cosmetic products are evaluated — Cosmetic/household product scores (0–100) are based on the highest-risk ingredient present; scientific sources include SCCS, ECHA, US EPA, AICIS, ANSES, and IARC
- Yuka Help Center — Cosmetic scoring category — Overview of how Yuka's cosmetic and household product scoring methodology works
- Think Dirty — Methodology — Dirty Meter scores products 0–10 across carcinogenicity, developmental/reproductive toxicity, and allergenicity/immunotoxicity; data sources include Health Canada, FDA, NTP, CDC, EPA, California Safe Cosmetics Program, Cosmetic Ingredient Review, EU Health Commission, and the ChemSec SIN List; undisclosed 'fragrance' is automatically scored 7 or higher
- EWG's Guide to Healthy Cleaning — Methodology — More than 2,000 household cleaning products graded A–F; grading pools data from 15 government, academic, and industry hazard/toxicity databases; ingredients with no hazard data default to a C grade rather than a pass
- U.S. FDA — 'Trade Secret' Ingredients (cosmetics labeling) — Fragrance ingredients can legally be listed simply as 'fragrance' without disclosing the individual chemicals
- California OEHHA — Proposition 65 list of chemicals — State list of chemicals known to cause cancer or reproductive harm, used as one hazard-data input by ingredient-scoring apps
Frequently asked
Which app is most accurate: Yuka, Think Dirty, or Bobby Approved?
None of them is running a lab test on your bottle — all three read the ingredient label and score it against public hazard databases, so 'accuracy' really means 'how good is the underlying data and the scoring rule.' EWG's Guide to Healthy Cleaning has the broadest data pool (15 pooled databases) and the most conservative default (unrated ingredients get a C, not a pass). Yuka is transparent that it scores off the single highest-risk ingredient, not an average. Think Dirty is the most aggressive about penalizing undisclosed 'fragrance.' Cross-checking two of the three on a borderline product is more reliable than trusting any single score.
Why do Yuka, Think Dirty, and EWG sometimes give the same product different scores?
They're not using identical databases or identical scoring math. Yuka leans on European bodies (SCCS, ECHA, ANSES) and scores off the worst single ingredient. Think Dirty pulls from about nine North American and European sources and auto-penalizes undisclosed fragrance. EWG pools 15 databases and grades on both hazard evidence and how much data exists at all. Different inputs and different formulas will produce different outputs for the same bottle — that's not a bug, it's three different methodologies looking at the same label.
Does an app give a bad score just because an ingredient hasn't been studied?
For EWG specifically, yes by design — an ingredient with no hazard data defaults to a C, not an A, precisely because 'nobody has studied it' is not the same claim as 'it's been studied and found safe.' That's a deliberately conservative choice. Yuka and Think Dirty handle unrated ingredients somewhat differently, but the underlying problem is the same across all three apps: most of the roughly 85,000+ chemicals in commercial use have limited public toxicity data, so every scanner is working with real, permanent gaps.
Do these apps actually test the product, or just read the label?
They read the label. None of the three apps sends your bottle to a lab. Each ingredient name is matched against a pre-built database of hazard assessments, and the product's score is calculated from that match. This means the score is only as good as what's printed on the bottle — which matters a lot, because in the US a manufacturer can legally list a mystery blend as simply 'fragrance' without naming a single compound in it.
What should I actually do with a low score from one of these apps?
Treat it as a lead worth investigating, not a verdict to panic over. Look at which specific ingredient triggered the score, then decide if it's a documented concern (an EPA, IARC, or Prop 65-listed hazard) or a low-confidence 'no data available' flag. If two of the three apps flag the same ingredient independently, that's a stronger signal than one app flagging it alone.