Anthropic and OpenEvidence Take Clinical AI to 100 Countries
A U.S.-built clinical decision-support platform is expanding free access across about 100 lower-income countries, testing whether local adaptation, evidence standards and governance can keep pace with global scale.
Physicians in about 100 low- and middle-income countries are set to receive free access to a clinical decision-support system built by Miami-based OpenEvidence and powered by Anthropic’s artificial intelligence. The companies announced the expansion during the United Nations General Assembly in New York, framing it as a way to narrow gaps in access to medical literature, specialist knowledge and continuing education. The initial country list includes Uganda, Angola, Sudan, Haiti and Mongolia, according to Reuters. The scale is striking, but the harder question is whether a tool trained and developed largely in wealthy markets can become clinically useful—and safe—when disease patterns, diagnostics, medicines, languages and referral options differ by country.
A U.S. clinical platform turns outward
OpenEvidence describes its product as a physician-facing search and synthesis layer over peer-reviewed research and treatment guidelines. The new program does not position the system as an autonomous diagnostician. Instead, clinicians ask questions, receive a synthesized answer with citations and retain responsibility for the decision. Anthropic will supply the underlying model infrastructure while OpenEvidence adapts the service to local settings. The companies disclosed no financial terms. In its own announcement, OpenEvidence said access would be free for physicians in participating lower-income countries. That distinction matters: removing a subscription cost can widen access to information, but it does not by itself provide connectivity, training, institutional oversight or confidence that recommendations fit local practice.
The program is regionally significant for the Americas even though much of its immediate impact will occur abroad. OpenEvidence is headquartered in Florida, maintains a San Francisco office, and launched the initiative with another U.S. technology company at a global forum in New York. Haiti is among the named countries, and the same platform is moving deeper into American clinical workflows. The University of Texas Medical Branch announced this week that it would integrate OpenEvidence inside its electronic health record so clinicians can reach cited medical evidence without leaving the care workflow. That integration shows the company pursuing two tracks at once: workflow adoption in U.S. health systems and broader access in lower-resource settings.
Localization is the core technical challenge
Clinical evidence is not portable in the same way as ordinary web search. A recommendation can be scientifically sound yet impractical where a laboratory test is unavailable, a medicine is not on the national formulary or the nearest specialist is hours away. Local epidemiology also changes the prior probability behind a diagnosis. OpenEvidence says its adaptations will account for infrastructure, available diagnostics and treatments, building on work in Rwanda and Botswana. A related partnership with Penn Medicine is deploying the platform to more than 10,000 clinicians and working through the Botswana-UPenn Partnership. The important feature is not simply exporting a model; it is involving clinicians who understand what care is actually possible.
A credible localization program therefore needs more than translated prompts. It requires maintained local guidelines, country-specific formularies, referral pathways, units, disease prevalence and escalation rules. It also needs a feedback process for recommendations that are technically correct but operationally unusable. PATH’s learning agenda for AI-enabled clinical decision support identifies six linked priorities: localization and language equity; real-world evaluation; voice and multimodal tools; local capacity and ownership; governance and trust; and enabling infrastructure. Those priorities are a practical checklist for this rollout. Smartphone availability may solve the last-mile interface in some settings, but it cannot substitute for governance, local evidence stewardship or a pathway to correct systematic errors.
The evidence supports caution, not dismissal
The best recent evidence from a lower-resource clinical setting is informative precisely because it is mixed. A pragmatic, cluster-randomized trial enrolled 9,691 patients cared for by 103 clinical officers at 16 primary care facilities in Kenya. Treatment failure within 14 days occurred in 2.2% of patients in the AI-assisted group and 2.0% in the control group; the adjusted comparison was not statistically significant. Investigators found no intervention-related serious adverse events or safety signal. Documentation quality improved, and the mean model cost was about four cents per patient, but the tool did not demonstrate a reduction in the primary clinical outcome. That is evidence of feasibility and some process benefit—not proof that clinical AI improves patient outcomes.
The broader literature is similarly unsettled. A 2026 evidence map reviewed 55 published studies and registered trials of large-language-model interventions in clinical practice. Human-AI collaboration dominated, but the authors found heterogeneous outcomes, inconsistent effects on efficiency and suboptimal reporting. Diagnostic accuracy in randomized studies ranged from 0.65 to 0.88, lower and more variable than in many nonrandomized studies. Physiological and biomarker outcomes were rare, while long-term follow-up was scarce. The implication is not that the technology lacks value. It is that benchmark scores, citation checks and usage totals cannot stand in for prospective evaluation of clinical decisions, patient outcomes, safety and equity in each deployment context.
Governance must travel with the software
The World Health Organization’s guidance warns that generative systems can produce false, biased or incomplete statements and can encourage automation bias, in which clinicians overlook errors because a machine produced the recommendation. WHO calls for stakeholder involvement, privacy protections, regulatory oversight and independent post-release audits when systems are deployed at scale. For a 100-country program, those principles translate into concrete operational questions: Who can see the queries? Are patient details retained? Which model version produced an answer? Can an institution reconstruct the evidence shown at the moment of care? How are harmful patterns reported across languages, and who has authority to suspend a flawed recommendation?
The answers may vary because clinical decision-support regulation is jurisdiction-specific. In the United States, the Food and Drug Administration’s January 2026 guidance explains which professional-facing software functions may fall outside the statutory definition of a medical device and which remain subject to device policies. Other countries have different regulators, data-protection rules and health-system accountabilities. Free access does not eliminate those obligations. A safe rollout needs clear intended-use boundaries, visible source citations, human review, documented escalation for uncertainty, security controls and local monitoring that can detect whether performance differs by language, geography or patient group.
Health-data architecture will be as important as model quality. A clinician’s question can itself reveal sensitive information even when a name is omitted, particularly in a small community or for a rare condition. Deployments should minimize patient-level data, separate clinical records from model-improvement pipelines, define retention periods and give institutions usable audit logs. Evidence provenance also needs version control: a citation visible today may be updated, withdrawn or superseded, while the model and retrieval system may change independently. Without a record of the query, source set, model version and final response, reviewers cannot reliably investigate a disputed recommendation. Those controls are not administrative extras. They determine whether a health system can learn from an error, compare performance over time and preserve professional accountability while still gaining the speed of an AI-assisted search.
What success should look like
The most meaningful success measures will be less dramatic than the country count. Health systems should track whether clinicians can access the service reliably, whether cited sources match local guidelines, how often recommendations are accepted or rejected, whether the tool changes prescribing or referral patterns, and whether any gains persist across rural and urban facilities. Evaluation should include patient-centered outcomes and harms, not only faster searches or more complete notes. Local institutions should participate in governance and publication of results rather than serving only as deployment sites. Public reporting of model changes, error categories and corrective actions would make the program more auditable and help other health systems judge whether the benefits transfer.
OpenEvidence and Anthropic are attempting to turn a U.S.-built clinical knowledge product into shared infrastructure for physicians who have had limited access to expensive evidence tools. That ambition is consequential and potentially useful. The evidence so far supports clinician augmentation, careful localization and active oversight—not deference to a model. If the partners pair free access with transparent evaluation, local ownership and durable safeguards, the initiative could narrow an information gap. If they measure success mainly by registrations and query volume, it may expand the reach of clinical AI faster than the knowledge needed to govern it.