The argument over animal research is often presented as a referendum on two competing visions of science: older experiments in living animals on one side, and newer technologies such as organoids, tissue chips and computer models on the other.
That is not yet the decision facing researchers or regulators. The more difficult question is narrower: whether a particular method can produce reliable information for a particular scientific or regulatory purpose.
That distinction matters because the United States is now making a serious institutional investment in human-based biomedical research. On September 21, the National Institutes of Health announced more than $88 million in infrastructure funding, a challenge program offering more than $7 million in awards, a new laboratory built around standardized human organoids, expanded recruitment of reviewers with expertise in human-based science and a request for information about publicly tracking vertebrate-animal use in NIH-funded research.
That same day, the Food and Drug Administration issued a direct final rule clarifying that scientifically valid non-animal methods may be used in the nonclinical development of drugs and biological products. The rule recognizes human-cell systems, organs-on-chips, computer models and other approaches. It does not require a particular method, and it does not lower the agency’s standards for evidence.
The announcements are significant. They are not proof that animal research has become obsolete.
A method is not validated merely because it is human
New approach methodologies, commonly called NAMs, include a large family of tools: organoids, tissue chips, cell-based assays, computational models and combinations of these approaches. They do not perform the same functions.
A liver organoid may help researchers study toxicity or metabolism. A tissue chip may reproduce selected interactions among cells under controlled conditions. A computational model may identify patterns in existing data or estimate how a compound behaves. None of those descriptions tells us, by itself, whether the method can predict the outcome needed for a clinical or regulatory decision.
Human origin is relevant. It is not a guarantee of validity.
Organoids and tissue chips simplify biology by design. They may omit the interactions among the nervous, immune, endocrine, circulatory and metabolic systems that shape the effect of a drug in a living organism. They may also vary according to cell source, culture medium, scaffolding, maturation and measurement protocol.
Those limitations do not make the methods useless. They define the questions for which they may be useful.
The National Academies’ assessment of new approach methodologies describes a pathway based on several related ideas: a clearly defined context of use, human biological relevance, technical reliability, reproducibility and fitness for the decision at hand. A method might be reliable for detecting one type of liver toxicity but not for predicting long-term immune effects. A model that helps explain a disease mechanism may not yet be adequate to determine whether a drug should enter a human trial.
The regulatory door is open, not empty
The FDA’s September rule changes the language surrounding nonclinical testing. It removes the implication that nonclinical evidence must necessarily come from animal studies and formalizes the possibility of using other scientifically appropriate methods.
That is an institutional change, but it is a permissive one. The agency is not declaring that every organoid, tissue chip or computational model is ready for approval decisions. It is saying that evidence need not be excluded because it was produced by a non-animal method.
The FDA’s March 2026 draft guidance on NAMs makes the same point in more operational terms. Developers are expected to explain what a method is intended to do, demonstrate that it is biologically relevant, establish technical reliability and show that it is fit for the regulatory purpose.
This is less dramatic than announcing the end of animal testing. It is more useful.
Regulators do not merely need interesting data. They need evidence that can support decisions involving human safety. That requires standards that allow reviewers to distinguish a promising research tool from a method whose performance has been demonstrated across laboratories, compounds and relevant biological conditions.
What the new NIH investment can—and cannot—show
NIH’s initiative recognizes that scientific methods are embedded in institutions. Researchers need equipment, protocols, training, reviewers and shared standards. A technology can be scientifically promising and still remain marginal if grant panels do not know how to evaluate it, universities cannot afford the infrastructure or regulators lack a framework for using the results.
The agency’s decision to recruit reviewers with expertise in human-based science therefore matters beyond personnel. It acknowledges that changing the model system also changes the people and institutions that decide what counts as credible evidence.
That is a different problem from the one WBC examined in its recent reporting on NIH’s shortened grant-review process. The earlier question was how the agency evaluates proposals under pressure from an application backlog. The current question is what kinds of evidence NIH is building the capacity to evaluate in the first place.
The new infrastructure may eventually make human-based research more reproducible and accessible. It may also expose a familiar problem in biomedical science: advanced tools can become concentrated in institutions with the money and personnel to operate them.
If sophisticated platforms are available only at a small number of elite centers, the transition could reduce animal use in some laboratories while widening inequalities in who can compete for grants, produce validation data and influence regulatory standards. NIH’s investment in infrastructure will therefore need to be judged not only by the number of facilities created, but by who can use them and whether the resulting methods are portable across laboratories.
The strongest case is not universal replacement
The evidence supports a narrower conclusion than the most enthusiastic public claims.
Human-based methods can provide useful information in particular contexts. Some animal models are poorly matched to particular human questions. Regulators are creating pathways for considering NAMs. And reducing animal use where a scientifically adequate alternative exists is both an ethical and a scientific objective.
What the evidence does not establish is that NAMs, as a category, are ready to replace animal research broadly. Nor does it show that fewer animal studies will automatically make drug development faster, cheaper or safer in every application.
The National Academies has explicitly warned that full replacement is not currently feasible for questions requiring complete multiorgan interactions and integrated biology. Its conclusion is not that alternatives have no value. It is that their value must be demonstrated for defined questions rather than assumed from their novelty or human origin.
This is why combination approaches may be more important than a single substitute. NIH’s challenge program is aimed in part at improving the measurement of biological signals in NAMs systems, including organ-on-a-chip devices. In practice, researchers may use several human-based methods together, or combine them with selected animal studies, clinical data and computational analysis.
That approach is less satisfying to people seeking a clean technological break. Biology rarely provides one.
The missing evidence is comparative evidence
The most important future studies will not simply show that an organoid can produce a result. They will compare the method prospectively with other models and with human outcomes.
Researchers should be able to ask: When a human-based method and an animal study disagree, which one predicts the clinical result? How often does each method produce false positives or false negatives? Are performance standards established before the result is known? Do findings replicate across laboratories with different staff, equipment and protocols?
Those questions are difficult partly because successful examples are easier to publicize than failures. A platform developer has an obvious incentive to emphasize cases in which a method worked. A credible validation record must also include ambiguous results, failed predictions and situations in which the model was not suitable.
The same standard applies to animal research. The existence of limitations in NAMs does not prove that animal models are reliably predictive. The relevant comparison is not between an idealized animal study and a flawed new technology. It is between the available methods for a specific decision.
Patients and animals both have a stake in the threshold
A premature transition could create false confidence in systems that omit important physiology, miss toxicities or produce inconsistent results. It could also move research capacity toward institutions able to afford complex platforms while leaving smaller laboratories dependent on standards they had little role in developing.
A transition that is too slow carries costs of its own. It may preserve experiments that add little useful information, delay human-relevant disease modeling and expose animals to procedures that could have been replaced or reduced.
The ethical argument for reducing animal use is therefore compatible with demanding evidence. In fact, it requires demanding evidence. An alternative that cannot answer the relevant question safely is not an adequate replacement merely because it avoids animal use.
The responsible question is not whether organoids are better than animals in the abstract. It is whether, for this particular decision, a specified human-based method produces more reliable information than the available alternatives.
NIH and FDA are building the infrastructure and regulatory pathways needed to answer that question. The transition has begun. The replacement problem has not been solved.













