Identify the promise
Separate the actual claim from the language, imagery, and proof points helping it land.
The North Test
We examine the claim, compare credible alternatives, and ask what the difference means in ordinary life.
See what we’re testingThe investigation
By the time a product reaches you, the conclusion has usually arrived first. The headline, comparison, proof point, and aesthetic have already been arranged.
The North Test backs up. We turn the pitch into a testable question, audit the evidence, choose the comparison that actually matters, and see what survives ordinary life.

What earns a North Test
A North Test begins when a product is asking you to believe something that could change what you buy, what you pay, or which version you choose—and there is a fair way to find out whether the promise holds up.
There needs to be a real question, a credible comparison, and enough uncertainty that the evidence could change our minds.
Otherwise, we are not testing anything. We are just confirming what we already thought.
How it works
Separate the actual claim from the language, imagery, and proof points helping it land.
Define what has to be true before the testing begins.
Use the alternative a buyer would actually face—not a convenient straw man.
Decide what will count before the result is known.
See what survives real use, then publish the verdict with the limits attached.
The Study Brief shows the question and constraints up front. The full methodology shows the rest.
Investigations now in design
Each question is locked before the answer is known. The question leads; the rigor is one step away. Open the Study Brief for the quick design, then the full methodology for every constraint, control, and limitation.
Study design publishedNorth Test 01 · Wine
Thirty reds from the same corner of Tuscany. One hundred twenty tasters. No labels, scores, or price cues. We begin with an already-considered $30 bottle, then ask whether Brunello—and each step up within it—actually buys more blind enjoyment.
A blind, randomized comparison designed to separate price, reputation, and label cues from what tasters actually enjoy.
The average 750ml table wine sold at U.S. retail is about $8.64. A $30 bottle is therefore not the cheap bottle—it is already a considered, premium-but-reachable purchase that many drinkers call “good.” That makes it the honest place to begin.
Latest U.S. retail scan data ↗Six Rosso di Montalcino wines around $25–$35 and 24 non-Riserva 2020 Brunellos, divided into lower, middle, and prestige price bands after a three-retailer price census. Same-producer Rosso and Brunello pairs are used where possible.
Brunello is not a grape. Both Brunello di Montalcino and Rosso di Montalcino are made from Sangiovese grown around Montalcino in Tuscany. Rosso is released younger; Brunello is aged longer and usually costs far more.
Official Montalcino wine definitions ↗Each taster receives eight randomized, three-digit-coded samples in identical glasses: two from each group. Temperature, pour size, and scoring conditions are standardized. No labels, critic scores, or prices appear until ratings are locked.
Overall liking is the primary outcome. Perceived quality, guessed price, willingness to pay, and “would choose again” are secondary. The Rosso-to-Brunello jump and the price-to-liking relationship within Brunello are analyzed separately.
Preregistration and independent statistical review are planned before testing. No tasting results are published yet. It can estimate patterns for these wines and these nonprofessional red-wine drinkers; it cannot declare a universal “best wine” or erase vintage, palate, and bottle variation.
Shortlist in progressNorth Test 02 · Red light
A handheld, a wearable mask, and a full-body panel solve different problems at radically different prices. Before choosing products, we are separating the claims, dose, and coverage from the fantasy that the biggest device must be the best.
Three device families, one declared outcome, and a specification audit before anything earns a place in the comparison.
The product shortlist is deliberately not locked. We first need one primary outcome—facial texture, pain, or recovery, for example—because no credible test can make every red-light promise its target at once.
Handhelds, hands-free face masks, and panels will be judged on what they actually cover, how long they take, and whether ordinary people will keep using them—not treated as interchangeable versions of the same thing.
Exact wavelengths, irradiance at the stated distance, delivered dose, schedule, supported indication, usable coverage, eye and heat precautions, return terms, and warranty all have to survive the specification audit.
The evidence and specification audit comes first. The shortlist will be published before purchase, followed by a standardized-use protocol with baseline measures, adherence, and a predeclared outcome.
Example randomized photobiomodulation trial ↗
Protocol in designNorth Test 03 · Luggage
A British boss taught Jude to read every dent in an aluminum Rimowa as proof of a life in motion. Her first checked case cracked on its first trip—to India. She kept buying the story anyway. Now the story gets tested.
Controlled stress plus real checked travel: durability, usability, repair access, and true cost per trip rather than patina alone.
Scratches and dents can signal a seasoned traveler. Cracked frames, missing wheels, fussy handles, and cases that no longer close are failures. The test separates appealing patina from damage that stops a bag doing its job.
Aluminum, polycarbonate, and soft-sided checked luggage will face the same loads and travel conditions, with carry-ons included where the use case is genuinely comparable. The brand shortlist is not yet locked.
Empty weight, usable capacity, wheel, handle, and closure cycles, rolling surfaces, impact, water intrusion, repair access, time without the bag, warranty exclusions, and true cost per trip.
Controlled stress testing will be paired with actual checked travel. Frequency matters: a bag used with children, a stroller, and regular baggage claim cannot be compared fairly with one that flies twice a year.
Rimowa’s official guarantee terms ↗Put a question on the table
Send the product, service, experience, or claim you cannot quite stop wondering about. The strongest questions are specific, consequential, and wrapped in a very good sales pitch.
The verdict standard
The claim is translated into something observable and testable.
Price, upkeep, time, repair, and the cost of changing your mind.
Not better in isolation—better than the credible alternative.
The trade-off, omitted condition, or inconvenience the pitch leaves out.
A useful verdict names the person and circumstance, not everyone.
The answer after ordinary use has replaced the first impression.
The method
A suitcase should not be tested like a skincare device, and a wine should not be tested like either. The protocol changes with the claim. The discipline does not: define the question before testing, choose a comparison that matters, separate signal from framing, disclose limitations, and show enough of the work for someone else to challenge the conclusion.
After completion
Every completed North Test publishes the verdict, who it is best for, the catch, the evidence that changed our mind—or did not—and the full methodology. If new evidence materially changes the answer, the verdict changes too.
See what earned a place