Phroneme
Validation architecture

A score is only as good as the design that defends it.

Each pillar below is a proposed validation requirement, not a completed method or reported result. It states the exposure a serious reviewer would attack and the design that Phroneme would need to implement and verify before publishing a CAI performance claim.

Proposed method · not yet implemented

The validation architecture describes proposed methods. Phroneme has not yet produced reliability, validity, fairness, normative, or longitudinal performance estimates, and no confirmatory study is preregistered.

Review status record →
01

The measurement model

The exposure

“Calibrated with IRT” is a slogan until you name the model. The wrong model on the wrong item format produces scores that look precise and mean nothing.

Proposed design answer

Proposed: Dichotomous items are calibrated with the two-parameter logistic model (discrimination and difficulty). Constructed-response and partial-credit items use Samejima’s Graded Response Model. Model fit is tested, not assumed.

  • Dimensionality is checked before calibration. Reasoning autonomy is treated as a hypothesis about structure, not a given; if the data are multidimensional, a bifactor or multidimensional IRT model is used rather than forcing a single score.
  • Local independence and item fit are evaluated for every item. Misfitting items are revised or retired before they ever reach a respondent.
  • The 3PL guessing parameter is admitted only where a pseudo-guessing floor is real, because it is unstable to estimate and is often a way to launder bad items.
02

The reliability of change

The exposure

The product is a trajectory, and the reliability of a change score is the quietest, deadliest problem in the field. A difference of two noisy measurements is noisier than either. Most longitudinal claims die here.

Proposed design answer

Proposed: We never report a bare difference. Each occasion carries a conditional standard error, and change is judged against a Reliable Change Index and a Minimum Detectable Change, then modeled as a latent slope rather than a subtraction.

  • Item response theory gives a standard error that varies with ability, so precision is reported where the person actually sits, not as a single test-level number.
  • A change is called real only when it exceeds the Minimum Detectable Change (about 1.96 × √2 × SEM) and the Reliable Change Index clears ±1.96. The on-site demo already enforces this.
  • The trajectory is estimated with a latent growth or latent change-score model, and individual slopes are shrunk toward the population (empirical Bayes) so a single lucky session does not masquerade as a trend.
  • The levers on change reliability are explicit: more information per occasion, optimal spacing between administrations, and pooling across occasions. We treat this as the central engineering problem, not a footnote.
03

Equated forms and exposure control

The exposure

A consequential score invites people to train for the test, not the capacity. A fixed form guarantees the score decouples from the construct. This is Goodhart’s law, and it is the default fate of the metric, not an edge case.

Proposed design answer

Proposed: A continuously refreshed bank delivers rotating forms, equated so scores are comparable, delivered adaptively so no two administrations match, with exposure controls and drift monitoring running underneath.

  • Forms are linked through a common-item nonequivalent-groups anchor design, placed on one scale by concurrent calibration or fixed-parameter linking with a Stocking–Lord transformation.
  • Item parameter drift is monitored continuously (robust-z and Lord’s statistic). An item whose behavior shifts is quarantined and retired before it distorts a trajectory.
  • Exposure is capped (Sympson–Hetter) and content is balanced across domains, so adaptivity cannot overfish a handful of items or narrow the blueprint.
04

Longitudinal measurement invariance

The exposure

You cannot interpret a trajectory across rotating forms unless the instrument measures the same construct the same way over time. Skip this and ordinary drift looks exactly like cognitive decline. This is the single most common invalid longitudinal claim.

Proposed design answer

Proposed: Invariance is established as a precondition, not an afterthought: configural, then metric, then scalar, then strict. Only the levels that hold are used to license comparison.

  • Metric invariance (equal loadings) is required before comparing associations over time; scalar invariance (equal thresholds) is required before comparing scores or slopes.
  • Where full invariance fails, partial-invariance models anchor on the items that hold, and the limitation is reported rather than hidden.
  • Anchor items in the equating design double as the invariance backbone, so the two systems reinforce each other.
05

Fairness and differential item functioning

The exposure

The moment a reasoning score touches a real decision it enters disparate-impact and accessibility law. A biased item is both an ethical failure and a legal one, and a license clause forbidding misuse does not survive a subpoena.

Proposed design answer

Proposed: Every item is screened for differential item functioning with two independent methods, flagged items are reviewed by content experts and removed, and the results are published rather than buried.

  • DIF is detected with Mantel–Haenszel and an IRT likelihood-ratio test, with purification passes, and classified by the ETS A/B/C effect-size scheme.
  • A flagged item is not quietly rescored; it is examined for the source of bias and removed if the bias is real. The audit trail is retained.
  • Fairness screening is also the legal shield described in the flagship paper: aggregate-first reporting starves the individual-harm predicate a disparate-impact claim requires.
06

Practice effects and norming

The exposure

Repeated testing raises scores on its own, independent of any real change, which directly confounds a trajectory instrument. And a number with no reference distribution cannot be interpreted at all.

Proposed design answer

Proposed: Retest effects are modeled explicitly as a fixed term in the growth model and suppressed by parallel equated forms; scores are placed against a stratified norming sample and always reported with a band.

  • The growth model carries a retest/exposure term so genuine change is estimated net of the practice artifact rather than contaminated by it.
  • Norms come from a demographically stratified reference sample; interpretation is percentile-with-interval, and a meaningful move is defined against the Minimum Detectable Change, not eyeballed.
  • Institutional reporting is aggregate and de-identified by default. The individual owns the score and can delete it.
07

The neural anchor, and where it stops

The exposure

EEG and fNIRS are seductive and are treated as ground truth by exactly the companies that get sued. Their own connectivity metrics have modest test-retest reliability, so anchoring a score to them imports their noise and their reverse-inference trap.

Proposed design answer

Proposed: The neural signal is never the score. It is convergent evidence and a difficult-to-clone credibility anchor, used on a consented subset, and its own unreliability is corrected for rather than ignored.

  • Behavioral CAI movement is tested against neural engagement change as convergent validity, disattenuated for the anchor’s own reliability so the correlation is not read too generously.
  • Neural data is the highest-sensitivity tier: minimized, firewalled, and governed by an explicit retention and destruction schedule, because biometric-privacy statutes punish collection mechanics, not only misuse.
  • Reverse inference bounds the anchor permanently. Activation corroborates behavior; it can never become the measurement. This is the design principle of the live brain.
Phase 0, stated to fail

What result would prove us wrong.

A construct that cannot say what would falsify it is not a measurement, it is a brand. The first study is not yet preregistered. These are proposed disconfirmable predictions and the results that would falsify them.

Test–retest stability

Equated forms yield acceptable short-interval reliability at the person level.

Falsified if

If parallel forms disagree beyond the modeled error, the bank is not equated and the trajectory is noise.

Convergent and discriminant validity

CAI correlates with established reasoning measures while separating from pure speed and vocabulary.

Falsified if

If it collapses into a speed or vocabulary factor, it is not measuring reasoning autonomy.

Sensitivity to offloading

Scores move in the predicted direction under experimental manipulation of cognitive offloading, net of practice.

Falsified if

If manipulation moves nothing once practice is modeled, the construct is not behaving as theorized.

Fairness

Differential item functioning is tested, reported, and used to remove biased items.

Falsified if

If DIF is pervasive and unremovable, the instrument is not deployable on real populations.

The honesty is not a concession. It is the moat.

Publishing the method, including its confidence intervals and its failure modes, is a scientific norm, a regulatory shield against the overclaim that has cost brain-training companies federal settlements, and a competitive position a marketing-led rival structurally cannot copy, because copying it would indict their own claims.