Phoenix APOE4 Research / Methods and limits

How to read Phoenix research

Phoenix research is longitudinal and observational. We follow real APOE4 carriers over time and compare each member with their own earlier self. There is no placebo group, no randomisation, and members chose their own interventions.

That means these are strong signals pointing at what to test next, not proof of cause. Most nutrition, supplement and lifestyle research you read online has exactly the same design. The difference is that we say so, and we show you the confounders below.

Documented limits
20
named plainly, grouped in four categories
Placebo arms
0
by design, every comparison is within-member instead
Independent re-analyses
Issue 01 was checked by analysts working blind to each other
What Phoenix research actually is

The largest ongoing member-led study of APOE4 carriers we know of.

Members log bloodwork, supplements and medications with start dates, daily adherence, daily check-ins, wearable sleep and heart-rate variability, and cognitive-game sessions. We link those streams together and look for what changed, when, and by how much.

The core comparison is within-member: a member against their own earlier self, or their own other days. That is deliberate. It removes every difference between people, age, genotype, sex, income, baseline health, diet, doctor, because both sides of the comparison are the same person. It is the single strongest thing about this design.

Three kinds of finding get published

Individual signals

n=1, n=2, n=3

One member, one timeline, one interesting change.

Group signals

Repeated

The same direction repeating across several eligible members.

Nulls

Published anyway

Things we expected to find and did not. We publish these too.

The full list of limits

20 limits, named plainly.

Grouped under design, measurement, statistics, and safety and scope. Each one is a real way this research can mislead you, and each is the reason we read every number below beside its caveat rather than alone.

Design

01

There is no placebo group.

Nobody in Phoenix was given a dummy pill. When a member's number improves, some of that improvement can be expectation, attention, and the simple act of paying attention to your health. Placebo effects on subjective ratings like energy and mood are large and well documented. They are smaller, but not zero, on objective measures like ApoB and wearable sleep duration.

02

Nothing is randomised.

Members chose what to take and when to start. That choice is not random, and the reason for the choice can be the real cause of the result.

03

Confounding by indication.

People often start an intervention because a number is already moving. Someone starts berberine because their lipids are climbing. Someone starts a sleep protocol during a bad month. Both make the intervention look worse than it is. The reverse also happens: someone starts a supplement during a motivated, everything-is-going-well stretch, which makes it look better than it is.

04

Co-intervention, and it is severe here.

Phoenix members stack. Members with repeat bloodwork log an average of 12 concurrent interventions; members with no bloodwork log 0.4. So “has repeat bloodwork” and “is a heavy stacker” describe nearly the same people, and any single-intervention signal could be the rest of the stack talking. For every individual case we publish, we state how many other interventions that member started nearby. When that number is zero, the case is much stronger. We say which.

05

Self-selection.

Phoenix members are not a random sample of APOE4 carriers. They are people who found out their genotype, sought out a community, and paid to join it. 59% carry two copies of APOE4, against about 2.7% in the general population, about a 22-fold enrichment. That makes Phoenix an extraordinary place to study APOE4 and a poor place to estimate what is typical.

06

Survivorship.

The members with the longest data histories are the ones who stayed. People for whom nothing worked are more likely to have stopped logging, which quietly tilts every long-window result positive.

Measurement

07

Daily check-ins are subjective.

Sleep quality, energy, mood, mental sharpness, calm and wellbeing are member-entered ratings on a 1–10 scale. They are real and they matter, but they are a person's impression of their day, not an instrument reading.

08

Some fields do not mean what their name says.

In the check-in data, mood and overall wellbeing are the same underlying field, so we never report them as two independent confirmations. The field named stress behaves as a calm score where higher is more relaxed. We publish these quirks rather than quietly working around them.

09

Lab noise can be bigger than the effect.

Every blood test has test-retest variability. HDL's noise floor is roughly ±7–8%, which is why we retracted a community-wide 4.6% HDL decline: the signal was smaller than the noise. Where a change sits inside the assay's own variability, we say so.

10

Units and naming are a minefield, and it has bitten us.

The same biomarker arrives from different labs under different names and in different units. ApoB and apolipoprotein_b are the same thing; LDL cholesterol, LDL particle number and LDL size are three different physical quantities that a careless match will average together into nonsense. Genotype arrives spelled four different ways. One naive match once discarded 866 of 1,277 members without any error appearing. We now canonicalise names, convert units explicitly, and print what we folded together.

11

Timing and lag are uncertain.

We know when a member logged a start. We do not always know when they actually started, whether they took it consistently, or at what dose. Dose is recorded on about 57% of entries and time of day is not captured at all. A follow-up blood test 77 days later is a real measurement at a real time, but the exposure between those two points is partly inferred.

12

Coverage is incomplete.

Not every member wears a device, retests on schedule, or logs every day. Missing data is rarely missing at random: people log more when things are going well.

Statistics

13

Multiple comparisons.

Test a thousand things and roughly fifty will look significant by chance alone. Issue 01 tested around a thousand habit-by-outcome comparisons plus ten biomarkers. Where we report a corrected result, we say which correction. Where we report an individual case, we are explicitly not claiming it survived correction. It is one interesting person, presented as one interesting person.

14

We slice the group to find the signal, and we tell you which slice.

When we report a pattern, we have often looked at a subgroup: one genotype, one sex, an age band, the members whose starting number was worst. Choosing the cut after seeing the data makes a finding easier to produce by chance, and we are not going to pretend otherwise. Our answer is not to stop doing it, because a signal in a subgroup is exactly what a carrier wants to know about. Our answer is that the cut is always named in the sentence. If a number describes twelve women over 55, it says so. If it describes everyone, it says that too. Read the cut before you read the number.

15

Small numbers.

Many of these cohorts are single or double digits. An n=4 signal where all four move the same way is genuinely interesting and genuinely fragile. Both things are true.

16

Regression to the mean.

People often start something when a number is at its worst. The next reading tends to be closer to their average regardless of what they did. This alone can manufacture an impressive-looking improvement.

17

Underpowered is not the same as ineffective.

Our sharpest test of this: we ran statins → LDL through the pipeline as a positive control. Statin users dropped 18.5 mg/dL more than non-users, which is the right direction and the right clinical magnitude for the most certain drug effect in cardiology, and it still landed at p = 0.154. If that cannot reach significance at this sample size, then a null result here means we cannot see it yet, not that the thing does not work. We hold ourselves to that reading.

Safety and scope

18

There is no adverse-event surveillance system.

Phoenix does not run structured safety monitoring. We cannot claim that no serious adverse events occurred among members, only that none were reported to us through the app.

19

Device and partner studies carry two extra limits.

Wearables are often not connected before a member's first use, so there is no clean group baseline to measure against. The comparison has to be use-nights against that member's own non-use nights instead. And when we report a device study, we report the discomfort too: in the ZenoWell taVNS study, alongside the positive sleep-onset reports, five members used pain language, two reported ear or skin irritation, two reported headache, and three used stop language. Positive experiences, null experiences and adverse experiences all get published.

20

This is not medical advice.

Phoenix does not diagnose, treat, cure or prevent any disease, including Alzheimer's. Medication decisions stay with your clinician. Phoenix does not check interactions with anything you already take. Nothing here is a recommendation to start, stop or change a treatment.

Not unique to Phoenix

This is true of most research you already read.

None of the above is unique to Phoenix. It describes the majority of the nutrition, supplement and lifestyle research that reaches the public, the studies behind the headlines, the podcast claims and the supplement labels. Most of it is observational. Most of it has no placebo arm. Most of it relies on people remembering what they ate.

The record of what happens when those observational findings finally get tested properly is sobering:

Observational findingWhat the randomised trial found
High beta-carotene intake tracked with LESS lung cancerATBC (1994), 29,000+ male smokers: 18% MORE lung cancer and 8% higher overall mortality in the beta-carotene arm. CARET (1996), 18,314 participants: 28% more lung cancer, 17% higher mortality. CARET was stopped early and participants were told to stop taking their vitamins.ATBC, NEJM 1994CARET, NEJM 1996
Vitamin E tracked with LESS heart diseaseHOPE (2000), 9,500+ high-risk patients over four years: no cardiovascular benefit. GISSI-Prevenzione, 11,000 heart-attack survivors over three-plus years: no preventive effect.HOPE, NEJM 2000GISSI-Prevenzione, Lancet 1999
Hormone therapy tracked with LESS coronary heart disease across large observational cohortsThe Women's Health Initiative (2002), 16,608 healthy postmenopausal women, stopped early at 5.2 years: more coronary events and more invasive breast cancer in the treated arm. The field then spent two decades reconciling the two, and the timing of initiation turned out to matter enormously.WHI, JAMA 2002

The lesson is not that observational research is worthless. It is that observational research is how you find the question, and a controlled trial is how you answer it. Phoenix is deliberately in the first business, and we are explicit about which business we are in.

What we will not do is what most of the sources our members already read do do: present an observational pattern with the confidence of a trial result, and leave the confounders out of the picture.

What Phoenix does about it

Disclosing limits is not the same as controlling for them.

These are the actual defences in the pipeline, all of which have caught real errors.

Within-member comparison, always. Every member is their own control group. It is not a placebo arm, but it eliminates every between-person confounder in one move.

The support-floor test. A real within-member effect gets stronger as you demand more days of evidence per member. Noise gets weaker. This one test killed a published headline (“stretching improves energy”) that had already survived multiple-comparison correction and an independent reproduction by a second analyst. We report which way every candidate moves.

Placebo tags as a control arm. We run mechanistically inert tags through the identical pipeline. When a vitamin with no plausible acute mechanism out-scored physical exercise, we knew the design was generating the signal, not the intervention. That finding was retracted.

Positive controls. We run a known drug effect through the pipeline to prove the method can detect something real before we trust it on something new.

Independent re-analysis. Issue 01 was analysed three times by analysts who had not seen each other's queries, then adjudicated line by line. Three bugs were found in our own analysis, and every one of them had pushed toward a more exciting answer. They always do. That is exactly what makes them invisible.

We publish the nulls. Issue 01's original headline died in review and we published its death. A screen of 263 supplement-biomarker pairs produced nothing that survived correction, and we published that too.

Reading guide

How to read a Phoenix number.

01

An individual case is one person. It tells you something interesting happened to somebody real, with dates and numbers attached. It does not tell you it will happen to you.

02

Read the co-start count. Zero means the member changed one thing. Eleven means you are looking at a whole protocol, and the honest claim is about the protocol.

03

Read the elapsed time. A 40% ApoB drop over 77 days and over 400 days are different stories.

04

Direction repeating across members is worth more than size in any one member. Four of four moving the same way beats one dramatic case.

05

When we say “we could not see it”, that is not “it does not work”. See limit 16.

Where this is going

What would make the next study stronger.

Every signal we publish comes with the test that would settle it. In general that means: a pre-specified retest date, a single intervention started alone, recorded dose and adherence, and where possible an alternating on-off schedule the member commits to in advance. That is the direction Phoenix is building, turning the interesting thing we noticed into the thing we planned to measure.

Read the signals next

Now see the signals these limits apply to.

Phoenix APOE4 Research publishes group signals, individual member cases, and the nulls that did not survive review, every one of them read beside the limits on this page.