Methodology
Last updated 22 September 2026 · 22 engines scored
Every score on this site is built the same way: collect what owners reported, keep only what independent sources agree on, and weigh it with the same rules for every engine.
We don’t review cars and we don’t score from spec sheets. Code does the counting and the arithmetic. Language models do the reading, and each one is asked one narrow question. None of them is told where the score bands fall.
Where the data comes from
Each report covers one engine in one car. An N47 in a 118d and an N47 in a 320d share a code, but they run at different temperatures, sit behind different gearboxes and fail differently, so they are researched and scored separately.
For each one we collect, strongest evidence first:
- Official recalls from the US regulator NHTSA and the Dutch vehicle authority RDW. A recall is the strongest evidence there is: the manufacturer has accepted the fault.
- NHTSA owner complaints, filed with the US regulator. Reported, but not confirmed by the manufacturer.
- UK MOT test records from the DVSA, covering 138 million vehicles. They show how far these engines really go and what they fail their annual test on. Testers never open the engine, so MOT data tells us about mileage and wear, not about what failed inside.
- Detailed owner histories, such as Bring a Trailer listings, which are dated and specific to one car.
- Owner forums for the make and model, such as Bimmerpost, E90Post, BimmerForums, GolfMK7, TDIClub, IH8MUD, Motor-Talk and PistonHeads, plus German, Russian and Polish boards and so on.
- Reddit posts and comments in car and brand communities.
Basic specs such as displacement, fuel type and production years come from Wikipedia and manufacturer data, checked against each other. Across the site that adds up to 1,419 sources behind 168 confirmed issues, and every engine report shows its own breakdown by platform.
Not in the score yet: manufacturer service bulletins, which are mostly behind paid workshop databases.
How reports become issues
Collect
Every source is searched for the engine code. Pages are split into individual posts, and each mention of a problem is kept as a short quote with the link it came from. Usernames never appear on the site.
Filter
A language model running on our own server checks that each quote is really about the part it mentions. A separate rule removes quotes about a different engine, which are common in comparison threads. Contradictions are thrown out automatically, such as an AdBlue fault reported on a petrol engine.
Group
Quotes are matched to a catalogue of 30 known failure types, from timing chain stretch to turbo failure, and grouped per engine.
Confirm
Each group gets a confidence score, explained below. Only confirmed issues appear on a report or count towards its score. Everything else stays a candidate until more evidence arrives.
Review
A person checks every engine before it is published: all the evidence that went into each decision, and why any evidence was left out. We also make sure the verdict and the final grade fit our rules and standards.
The model that reads the posts runs on our own server, not a paid cloud service. Reading a post costs us nothing, so there is no reason to skim or sample: it reads every one, however many there are, often hundreds for a single engine. It was tuned on a set of 50 engines whose problem lists and grades were made by people.
When an issue is confirmed
Confidence starts from the strongest single source behind an issue, out of 100:
- Official recall: 100
- NHTSA complaint: 70
- Detailed owner history: 35
- Forum post: 25
- Reddit post: 20
A report that gives a specific diagnosis, shows evidence or confirms that a repair fixed the problem earns up to 25 extra points. Each additional independent platform reporting the same issue adds 5, up to 30. Five posts in one thread count as one platform, so a single viral thread cannot manufacture a consensus.
An issue is confirmed at 40 points, or straight away if there is an official recall. That confidence number is the one shown next to every issue on an engine report.
Wear or design flaw
The same failure can be normal on one engine and a defect on another. A timing chain that needs attention at 300,000 km is wear; the same chain failing at 90,000 km is a design problem. So every confirmed issue is classified as one of six types:
- Expected maintenance. Parts meant to be replaced on a schedule. Light impact on the score.
- Age-related wear. Parts wearing out at a reasonable age or mileage. Light impact.
- Environmental. Caused by climate, fuel quality or how the car is used, such as only short trips. Light impact.
- Owner-caused. Neglect or modification. Counted, but weighted lightly, so one abused car does not define the engine.
- Design weakness. The design makes the part fail early even when the car is looked after. Heavy impact.
- Manufacturing defect. Faulty parts or assembly from the factory. The heaviest impact.
An official recall, or a documented factory revision of the faulty part, makes an issue a defect straight away. Otherwise it takes at least two kinds of evidence agreeing: official complaints, failures well before the part’s normal life, and consistent owner reports. Owner reports alone cannot make something a defect. A model also reads the owner quotes behind each issue to judge whether they describe a defect or normal wear. Its verdicts are only applied when most of an engine’s issues point the same way, so a good engine with plenty of small complaints is not buried under them.
Each confirmed issue is also rated on two separate questions:
- Severity, meaning how bad it is when it happens: the risk of catastrophic failure (35%), damage to other parts (25%), the chance of leaving you stranded (20%) and the repair cost (20%), minus 10 points if maintenance prevents it. The result is Critical (80+), High (60–79), Moderate (35–59) or Low.
- Frequency, meaning how often it happens, measured as a share of everything reported about that engine: very rare (under 2%), rare (2–5%), occasional (5–10%), common (10–20%) or very common (20% and over). An issue reported on only one platform can never rate above occasional.
How engines are graded
The final score averages three independent assessments, then maps the result onto a fixed scale.
1. The evidence score
Five questions, each answered in a separate call by a model that sees only the evidence relevant to that question. A servicing-cost question never sees catastrophic failures, for example. Each question is answered twice and averaged:
The weighted total is then stretched onto 0–100 by one fixed formula, the same for every engine.
2. Two reputation checks
Language models from two different companies are each asked, three times, for the engine’s established reputation among owners and mechanics. People post when something breaks, not when an engine is still fine at 300,000 km, so forum evidence on its own undersells good engines. The two models also make different mistakes, so together they are steadier than either one alone.
3. Calibration
The average of the three is mapped onto the final scale using 23 reference engines whose real-world records are well established, from famously durable to notorious. An engine cannot score beyond the best or worst of them, so scores run from about 18 to 95. The same bands then apply to every engine, so a 70 on a BMW means the same as a 70 on a Toyota:
The recommendation on each report is decided by code from the score and the confirmed issues. A model then writes the short verdict, the buying advice and the summary of what owners report. It explains numbers that already exist and is not allowed to add an issue or a figure of its own.
Repair costs
Each of the 30 failure types carries a typical repair price range in euros. Timing chain and guide work, for example, is €1,000–2,000. The same failure shows the same range on every engine, and the cost also counts for 20% of an issue’s severity. Treat it as a guide to the size of the bill, not a quote: it does not yet reflect engines where the same job means taking the engine out, or labour rates where you live.
Independence
- No industry money. We take nothing from manufacturers, dealers or parts sellers. There are no sponsored engines or paid placements, and nobody can pay to change a score.
- Ads don’t touch the data. The site is paid for by ads served through Google AdSense. Google chooses the ads, which can include car brands, and advertisers have no say over any score or report.
When scores change
Scores are not permanent. When we collect new evidence for an engine or revise the method, affected scores are recalculated with the same rules, and every report shows when it was last updated and which scoring version produced it. Changing a prompt, a model or a weight means re-measuring against the reference engines first: the pipeline refuses to score until that is done.
If you know something we have missed, use “Something wrong here?” at the foot of any engine report, or the contact form. Every message is read by a person. Mileage, symptoms and what the repair cost are the details that change a score.
What a score cannot tell you
- Owners post problems.Happy owners rarely write in. The reputation checks balance that, which means part of every score reflects an engine’s established reputation, not only the reports we collected.
- Rare engines have thin evidence. A score built on a handful of reports is less certain than one built on hundreds. Every issue shows how many sources back it and how confident we are in it.
- It was tuned on a limited set of engines. The reading model on 50 engines graded by people, the final scale on 23 reference engines. Checking the method against engines it has never seen is the next step.
- It is not your car. A score describes an engine across many owners, not the condition of the one in front of you. Poor maintenance can ruin a good engine and good maintenance can save a weak one, so always have the car inspected before you buy it.