GalenOps
Background
GalenOps is a rudimentary implementation of the NxtCure AI-CRO vision: automate the repeatable, document-heavy Contract Research Organization functions in phased horizons, and build the company around a cross-functional agent orchestration layer that coordinates every module end to end. The repository lives at ~/void/www/git/medicalapps/galenops and is not source controlled in this monorepo.
A rudimentary implementation of the NxtCure AI-CRO vision: automate the repeatable, document-heavy CRO functions in phased horizons, keep a human-in-the-loop as a feature, not a limitation, and build the company around the one durable differentiator — a cross-functional agent orchestration layer that coordinates every module end to end and escalates to a human only on exception.
--GalenOps README
What is an agent? UNIX + LLM. GalenOps takes the reductionist reading seriously: the “AI” in the product is deterministic Elixir with explicit seams marked for models, so trial management gets guarantees first and inference later. Every automated judgment that is medical, safety, regulatory, or financial in nature terminates in a human review queue rather than a side effect.
Applications
The repository is a four-app Phoenix 1.8 monorepo. There is no umbrella; four independent Mix projects are wired together over HTTP:
galenops_backend(port 4000): Phoenix JSON API with Ecto over SQLite (ecto_sqlite3). All domain logic, all state, the automation modules, the orchestration layer, and the review queue.galenops_portal(port 4010): LiveView sponsor/CRO operations console styled with daisyUI.galenops_mobile(port 4020): mobile-first LiveView PWA for field workflows — review queue, safety triage, site monitoring. A drafted LiveView Native SwiftUI host sits innative/but is blocked onlive_view_native 0.4.0-rc.1pinning Phoenix~> 1.7.galenops_luna(port 4030): Luna, the standalone patient companion app ported from the NxtCure mobile app — chat, structured check-ins, a medication and symptom tracker with a scoreboard, and optional voice.
Only the backend has a Repo. The portal, mobile, and Luna apps hold no state of their own; each is a pure client of the backend JSON API through a thin Req wrapper. The portal version reads as:
defmodule GalenopsPortal.Backend do
@moduledoc """
Thin HTTP client for the Galenops backend JSON API.
All portal LiveViews go through this module; the portal holds no state of
its own. Functions return plain decoded maps; `{:error, reason}` on
transport or non-2xx responses.
"""
def dashboard, do: get("/dashboard")
def studies, do: get_data("/studies")
def study(id), do: get_data("/studies/#{id}")
def create_study(attrs), do: post("/studies", %{study: attrs})With the HTTP plumbing at the bottom of the module amounting to little more than:
defp get(path, params \\ []) do
handle(Req.get(base_url() <> path, params: params, retry: false))
end
defp handle({:ok, %Req.Response{status: status, body: body}}) when status in 200..299,
do: {:ok, body}
defp handle({:ok, %Req.Response{status: status, body: body}}),
do: {:error, "backend returned #{status}: #{inspect(body)}"}
defp handle({:error, exception}),
do: {:error, "backend unreachable: #{Exception.message(exception)}"}This is the entire integration surface between the four services. No shared database, no message bus, no service mesh. Well-ordered filesystem, efficient automation.
Domain Model
The backend organizes its Ecto contexts into the four pillars of CRO work, in galenops_backend/lib/galenops/:
Galenops.Studies(Pillar 1, study start-up and design): studies, versioned protocols, feasibility assessments, regulatory submissions, and site contracts.Galenops.ClinicalOps(Pillar 2, execution): sites with a KRIrisk_score, pseudonymous patients keyed byexternal_ref, monitoring visits, and supply forecasts.Galenops.DataSafety(Pillar 3, data and safety): EDC data queries and adverse events.Galenops.Quality(Pillar 4, documentation and quality): TMF/CSR/TLF documents, QA findings, and site payments.
Two more contexts carry the differentiator: Galenops.HumanLoop is the review queue, and Galenops.Orchestration is the agent-run audit trail plus a twelve-step study lifecycle checklist. Finally Galenops.Luna treats the participant companion as a retention and safety sensor: chat messages with a classified intent, check-ins rating feeling 1 through 5, tracked items, item logs, and appointments. All schemas use binary_id primary keys and UTC timestamps.
Automation Engine
Galenops.Automation is a single 1089 line module of pure functions over the domain: draft_protocol/1, assess_feasibility/2, forecast_supply/2, score_site_risk/1, triage_adverse_event/1, resolve_data_query/1, draft_csr/1, generate_data_queries/1, generate_tlfs/1, audit_tmf/1, check_site_activation/2, predict_retention/1, and watch_termination_risk/1. The heuristics are not hand waving: the module carries the SPIRIT 2013/2025 protocol completeness checklist as @spirit_checklist, a fourteen day query-aging SLA as @query_aging_sla_days, and a kit buffer of 1.15 for supply forecasting. Each red flag rule in watch_termination_risk/1 is traceable to a citation in docs/intent/lit_review_synthesis.md.
The escalation posture is visible directly in the code. Adverse event triage never lets software decide causality:
def triage_adverse_event(%AdverseEvent{} = ae) do
if ae.serious or ae.severity == "severe" do
{:ok, ae} = DataSafety.update_adverse_event(ae, %{status: "under_review"})
{:ok, task} =
HumanLoop.open_task(
"adverse_event",
ae.id,
"assess_serious_event",
"Serious/severe adverse event requires safety physician assessment: #{String.slice(ae.description || "", 0, 140)}",
"critical"
)
{:ok, %{adverse_event: ae, review_task: task}}
else
{:ok, ae} =
DataSafety.update_adverse_event(ae, %{
status: "assessed",
assessment:
"Auto-assessed as non-serious (#{ae.severity}). Included in aggregate safety reporting; no expedited report required."
})
{:ok, %{adverse_event: ae, review_task: nil}}
end
endAlongside the automation engine, Galenops.Integrations marks the build-versus-partner boundary with three simulated vendor surfaces: MatchingEngine for EHR/RWD candidate matching, SupplyVendor for IRT/RTSM forecasting, and PaymentsVendor for milestone site payments. The matching engine never auto-enrolls; the recruitment agent always opens an import_matched_candidates review task.
Agent Orchestration
Galenops.Orchestration.Pipeline in lib/galenops/orchestration/pipeline.ex is the coordinator. An ordered list of thirteen agents walks the study lifecycle: feasibility_agent, protocol_agent, site_activation_agent, recruitment_agent, monitoring_agent, cdm_agent, safety_agent, retention_agent, termination_watch_agent, supply_agent, biostats_agent, medical_writing_agent, and tmf_agent. Each agent is a defp run_agent(:name, study) clause returning {:done, reasoning, task_or_nil} or {:skip, reasoning}, calling into Galenops.Automation and Galenops.Integrations.
The main loop wraps every agent so that a crash is itself an escalation, never a silent failure:
def run_study_pipeline(%Study{} = study) do
runs = for {agent, horizon} <- @steps, do: execute(agent, horizon, study)
{:ok, runs}
end
defp execute(agent, horizon, study) do
{outcome, reasoning, task} =
try do
case run_agent(agent, study) do
{:done, reasoning, nil} -> {"completed", reasoning, nil}
{:done, reasoning, task} -> {"escalated", reasoning, task}
{:skip, reasoning} -> {"skipped", reasoning, nil}
end
rescue
exception ->
{:ok, task} =
HumanLoop.open_task(
"study",
study.id,
"resolve_agent_exception",
"Agent #{agent} crashed on \"#{study.title}\": #{Exception.message(exception)}. Orchestration continued; this step needs human follow-up.",
"high"
)
{"failed", "Unhandled exception: #{Exception.message(exception)}", task}
endEvery execution is persisted as an AgentRun with a plain-language reasoning string, an outcome of completed, escalated, skipped, or failed, and a pointer to the review task when one was opened.
The safety agent illustrates the do/skip/escalate shape and closes with a sentence worth keeping:
_ ->
{:done,
"Triaged #{length(reported)} adverse event(s); #{length(escalated)} serious/severe escalated to the safety physician. Causality is never decided by software.",
task}Galenops.Orchestration.checklist/1 then derives the portal-facing twelve step checklist by joining the latest AgentRun per agent against its ReviewTask, yielding not_started, attention, awaiting_review, or complete per step.
Human In The Loop
Galenops.HumanLoop in lib/galenops/human_loop.ex is the other half of the contract. Agents open tasks through one function, deduplicated so pipeline re-runs do not stack duplicate escalations:
def pending_task_exists?(subject_type, subject_id, action) do
Repo.exists?(
from t in ReviewTask,
where:
t.subject_type == ^subject_type and t.subject_id == ^subject_id and
t.action == ^action and t.status == "pending"
)
end
def open_task(subject_type, subject_id, action, summary, risk_level) do
create_review_task(%{
subject_type: subject_type,
subject_id: subject_id,
action: action,
summary: summary,
risk_level: risk_level,
status: "pending"
})
endApproval and rejection side effects are keyed on the ReviewTask.action string: an approved approve_protocol_draft marks the protocol approved, a rejection returns it to draft, and a serious adverse event is never auto-closed on rejection. Known actions include approve_protocol_draft, approve_csr_draft, assess_serious_event, resolve_data_query, submit_regulatory, release_site_payment, validate_tlf_outputs, confirm_onsite_visit, review_low_feasibility, import_matched_candidates, assess_termination_risk, resolve_agent_exception, and the three Luna escalations luna_crisis_escalation, luna_severe_symptom, and luna_withdrawal_signal. A companion module HumanLoop.TaskContext polymorphically assembles details, consequences, agent reasoning, and links per subject type so the reviewer can decide well from one screen.
Luna
Luna is the participant side, ported from the NxtCure app. Galenops.Luna.Brain is where an LLM would be: the NxtCure chat system prompt has been ported into deterministic rules — ordered keyword precedence over crisis, severe symptom, withdrawal, medication, visit, and status phrase lists — so the role contract (never a doctor, crisis resources first, one-sentence deflection out of scope) is enforced in code rather than in a prompt. An LLM can be slotted behind the same intents later.
The one live external AI dependency is Deepgram, voice only and optional. Galenops.Luna.Deepgram makes two plain Req calls: speech to text against /v1/listen with nova-3, and text to speech against /v1/speak with aura-2-aurora-en. Without a DEEPGRAM_API_KEY the voice endpoint reports disabled and text chat is unaffected. A spoken turn is one round trip: audio in, transcribe, route through the same Luna.chat/2 path as typed messages, speak the reply, return base64 audio.
In the LiveView, escalation surfaces to the participant as a flash while the review task lands in the CRO queue in real time:
case Backend.chat(user["id"], text) do
{:ok, %{"reply" => reply, "escalated" => escalated}} ->
sent = %{"role" => "user", "content" => text, "escalated" => escalated}
socket =
socket
|> assign(transcript: socket.assigns.transcript ++ [sent, reply], error: nil)
|> then(fn s ->
if escalated,
do: put_flash(s, :error, "Escalated to the study team — a human will follow up."),
else: s
end)
{:noreply, socket}API Surface
The router at galenops_backend/lib/galenops_web/router.ex exposes RESTful resources per pillar, plus the orchestration, review queue, Luna, and automation routes:
# Horizon 3 — cross-functional agent orchestration + audit trail
post "/studies/:id/orchestrate", OrchestrationController, :run
get "/studies/:id/checklist", OrchestrationController, :checklist
resources "/agent_runs", AgentRunController, except: [:new, :edit]
# Human-in-the-loop review queue (grouped route must precede :id routes)
get "/review_tasks/grouped", ReviewTaskController, :grouped
resources "/review_tasks", ReviewTaskController, except: [:new, :edit]
get "/review_tasks/:id/context", ReviewTaskController, :context
post "/review_tasks/:id/approve", ReviewTaskController, :approve
post "/review_tasks/:id/reject", ReviewTaskController, :reject
# Automation endpoints — each proposes work and routes risk to a human
post "/studies/:id/automation/draft_protocol", AutomationController, :draft_protocol
post "/studies/:id/automation/assess_feasibility", AutomationController, :assess_feasibility
post "/studies/:id/automation/forecast_supply", AutomationController, :forecast_supply
post "/studies/:id/automation/draft_csr", AutomationController, :draft_csrTwo smaller modules round out the backend: Galenops.PdfWriter is a dependency-free PDF renderer (Courier, US Letter) for protocol, CSR, and TLF previews served at GET /api/documents/:id/pdf, and Galenops.Literature provides pharmacovigilance mechanism evidence through NCBI E-utilities, honestly separated into always-available reproducible PubMed search links and best-effort live article lookups that degrade to empty on network failure. Never fabricated citations.
Demo Data and Seeding
The seed pipeline has three layers. Galenops.Release.maybe_seed/0 is the boot-time guard, wired into compose.yml as the backend command so first boot migrates, seeds if the database is empty, then serves:
def maybe_seed do
cond do
System.get_env("SEED_ON_BOOT", "false") not in ~w(true 1 yes) ->
IO.puts("SEED_ON_BOOT disabled — skipping seed.")
not empty?() ->
IO.puts("SEED_ON_BOOT: database already has studies — skipping seed.")
true ->
IO.puts("SEED_ON_BOOT: empty database — seeding demo data...")
seed()
end
endpriv/repo/seeds.exs creates the demo study GAL-001: Adaptive Phase II in Refractory Melanoma, deliberately behind enrollment pace so the termination watch agent raises an accrual flag, walks it through every automation function pillar by pillar, seeds three Luna participants with distinct narratives (a perfect-day streak, missed medications with withdrawal talk, symptom ratings), and finishes by running the full orchestration pipeline so the reasoning trail and review queue are populated on first page load.
Galenops.DemoImport then ingests real-scale CDISC SDTM extracts from priv/demo_data/*.json, produced by scripts/extract_sdtm_study.py (pandas plus pyarrow, run under uv): a synthetic Phase 3 NSCLC study with 450 subjects and 2,840 adverse events, and the CDISC pilot Alzheimer’s study with 306 subjects. The import mapping keeps the safety agent honest — the five most recent serious adverse events stay reported so a live triage queue exists, and major protocol deviations become open QA findings feeding the site KRIs.
Regulatory Readiness
We’re building the compliant AI workspace for regulated labs, made for GMP, GLP and 21 CFR Part 11 from day one … Every tool that touches a study has to be validated first, which is exactly why AI has passed them by.
--Ordo Labs launch post (YC F26)
That post names four regulatory items: GMP, GLP, 21 CFR Part 11, and validation of every tool that touches a study. This section is the plan for GalenOps to be able to make the same claim honestly, and to survive a customer’s vendor audit when it does. It is a working plan, not legal advice; the scoping decisions in Phase 0 should be confirmed with a regulatory/QA consultant before money is spent on Phases 3 and 4.
What can and cannot be obtained
None of the four is a certificate a software company is issued, and saying “GMP certified” or “Part 11 certified” about software is a red flag to any QA auditor. Each one works differently:
- GMP (21 CFR 210/211 in the US, EudraLex Volume 4 in the EU) and GLP (21 CFR Part 58, the OECD Principles of GLP) are predicate rules that bind the regulated facility — the manufacturer or the non-clinical test facility. Inspectors (FDA, or in Sweden the Medical Products Agency for GMP and Swedac as the GLP monitoring authority) inspect the lab, including the computerised systems it relies on. A vendor cannot be GMP or GLP compliant on its own behalf; it can supply a system and evidence that let the customer stay compliant.
- 21 CFR Part 11 governs electronic records and electronic signatures used to meet any predicate rule. It is met by technical controls in the product plus procedural controls at the customer. The vendor’s job is the first half and the documentation that proves it. EU Annex 11 (Computerised Systems) is the European counterpart and should be covered in the same work.
- Validation is performed by the regulated user for its intended use, but in practice it rests on the vendor’s own documented testing. The accepted frameworks are ISPE GAMP 5 (2nd edition) and FDA’s Computer Software Assurance (CSA) guidance, which favours risk-based, critical-thinking assurance over paperwork-heavy scripted testing.
GalenOps runs clinical trials rather than labs, so the predicate rule it actually lives under is GCP — ICH E6(R3), which added explicit expectations for computerised systems — not GMP or GLP. GMP and GLP matter only if we sell into manufacturing QC or non-clinical (preclinical) labs, which is Ordo Labs’ market. Part 11 and validation apply either way. So what GalenOps can obtain is:
- a Part 11 / Annex 11 compliance assessment of the product, clause by clause, with evidence;
- a validation package customers can adopt (GAMP 5 / CSA);
- a quality management system (QMS) that makes both repeatable and auditable;
- independent evidence: a successful customer audit or mock inspection, and, because every vendor questionnaire asks, a security attestation (SOC 2 Type II or ISO 27001).
Phase 0: Scope
- Decide the target customers for the next twelve months: CROs and sponsors (GCP only), or also preclinical labs (adds GLP) and manufacturing QC (adds GMP). Each predicate rule added means more SOP coverage and more intended-use validation.
- Write the intended use statement for GalenOps: what records it creates, which of them are GxP records, and which decisions it supports. This single page drives the risk assessment and every test afterwards.
- Classify the system under GAMP 5 (configurable product plus custom code) and record which records fall under Part 11 — in GalenOps at least protocols, adverse event assessments, data queries, CSR/TLF drafts, TMF documents and every
ReviewTaskdecision. - Exit: a signed scope memo, reviewed by an outside regulatory consultant.
Phase 1: Quality management system
Auditors read the QMS before they look at the software. The minimum set of SOPs for a GxP software vendor:
- document control and records retention;
- software development lifecycle (requirements, design review, code review, testing, release);
- change control and configuration management, including how a hotfix is released;
- risk management (ICH Q9-style, applied to software functions);
- deviation, incident and CAPA handling;
- training, with training records per person;
- supplier management — Deepgram, NCBI and any hosting provider are suppliers;
- backup, restore, business continuity and disaster recovery, with a tested restore;
- security incident response and access management.
ISO 9001 certification is optional, but it is a recognised external signal that the QMS runs. Exit: SOPs approved and in effect, with training records for everyone who touches the code.
Phase 2: Part 11 technical controls in GalenOps
GalenOps already has the right posture — every medical, safety, regulatory or financial judgement lands in a human review queue, and every agent execution is persisted as an AgentRun with its reasoning. What it does not yet have is a Part 11-grade record of who decided what and when. The gap list, by clause:
- §11.10(a) validation — covered by Phase 3.
- §11.10(b) accurate and complete copies — export of any record with its full audit trail, in human-readable (the existing
PdfWriter) and electronic form. - §11.10(c) protection and retention — records must be retrievable for the whole retention period (GCP: at least as long as the sponsor requires, often 25 years under the EU CTR). SQLite in a Docker volume needs a documented backup, restore and archival design; migrations that run on every boot must be proven not to alter existing records.
- §11.10(d) and (g) access and authority checks — named user accounts, roles (reviewer, safety physician, QA, admin) and enforcement that only the right role can approve a given
ReviewTask.action, such asassess_serious_event. - §11.10(e) audit trail — secure, computer-generated, time-stamped, recording who created, changed or deleted each GxP record, without obscuring the previous value. Today a review task’s
statusis updated in place; that needs an append-only audit table (old value, new value, user, UTC time, reason) written in the same transaction as every change, with no API or UI path to edit or delete it. - §11.10(f) operational checks — enforce step order where the process requires it (e.g. a protocol cannot be
approvedwithout a completed review). - §11.10(h), (i), (j), (k) — device checks where relevant, training evidence, a written accountability policy for signatures, and control of system documentation. These are mostly Phase 1 procedure.
- §11.50 and §11.70 signature manifestation and linking — an approval must display the signer’s printed name, the date and time, and the meaning (approval, review, responsibility), and the signature must be cryptographically or otherwise inseparably bound to the exact record version signed.
- §11.100, §11.200, §11.300 electronic signatures — unique per person, identity verified before issue, two components (user ID plus password) re-entered at signing, password ageing and lockout rules, and the one-time letter of certification to FDA that the organisation’s e-signatures are the legal equivalent of handwritten ones (filed by the regulated customer, but the product must support it).
- Data integrity (ALCOA+) — the FDA and MHRA data integrity guidances expect records to be attributable, legible, contemporaneous, original and accurate. The Luna and agent paths should record the source of every value (human, agent, integration).
- AI-specific — every agent output that becomes part of a record carries the agent name, code or model version and inputs used, so it can be reproduced. FDA’s January 2025 draft guidance on AI to support regulatory decision-making for drugs asks for exactly this credibility evidence, risk-scaled by the model’s influence on the decision. The deterministic Elixir engine is an advantage here; any LLM slotted in later must be version-pinned and covered by the same change control.
Exit: a clause-by-clause Part 11 / Annex 11 matrix for GalenOps, each row pointing to the code, test or SOP that satisfies it.
Phase 3: Validation package
Deliver a package a customer’s QA can adopt instead of writing their own:
- User and functional requirements specifications, derived from the intended use.
- A risk assessment per function (GAMP 5 / CSA): high-risk functions such as safety triage, approvals and audit trail get scripted testing; low-risk ones get unscripted or exploratory testing with recorded results.
- Test evidence: installation checks for the Docker deployment, operational tests (the existing automated tests count if they are traceable and their runs are retained), and performance qualification scenarios using the seeded GAL-001 study and the SDTM demo imports.
- A requirements traceability matrix and a validation summary report.
- Release notes and an impact assessment for every version, so customers can do change-control revalidation quickly.
Exit: the first release shipped with a complete package, and the package reviewed by an independent CSV consultant.
Phase 4: Independent evidence
- A mock vendor audit or mock inspection by an outside GxP auditor, with findings closed through CAPA.
- A SOC 2 Type II report or ISO 27001 certificate for the hosted service; almost every pharma vendor questionnaire asks for one before it asks about Part 11.
- The first paying customer’s vendor qualification audit, passed. This is the evidence that matters most, and the one Ordo Labs is pointing to with its paid pilots in Sweden.
- If Phase 0 chose GLP or GMP markets: a customer lab that has taken GalenOps through its own GLP or GMP inspection with no findings on the system.
Order and dependencies
Phase 0 comes first and is cheap. Phase 1 and Phase 2 run in parallel, because the audit trail and e-signature work is engineering while the SOPs are writing. Phase 3 needs Phase 2′s controls to exist before they can be tested. Phase 4 needs everything before it. The single most important engineering item is the append-only audit trail with signature linking on ReviewTask decisions: it is the part a Part 11 auditor tests first, and the one GalenOps’ human-in-the-loop design already points towards.
Deployment
Bring-up is one command from the repository root, with secrets in a gitignored .env generated via mix phx.gen.secret:
docker compose up -d --buildThe backend healthcheck curls /api/dashboard with a 600 second start period because first boot seeds before listening; the portal, mobile, and Luna services gate on condition: service_healthy. SQLite persists in a backend_data volume and migrations run on every boot. Local development is the standard mix setup && mix phx.server per app.
The docs/architecture/ directory carries six Graphviz control-flow diagrams (system context, request flow, orchestration, human loop, module map, boot), a Touying typst slide deck in slides.typ compiled to a 16:9 PDF, and a build.sh that renders the dot sources and compiles the deck. The literature review in docs/intent/lit_review_synthesis.md maps each implemented heuristic back to its citation.