📘
Technical Briefings · 44 min read · September 24, 2026
How veterinary AI can reduce claim denials
How AI can review a veterinary insurance claim before it is filed, cite its sources and leave every decision with the clinician.
By Quill Direct Vet Care

A technical briefing in six parts from Quill Direct Vet Care, covering the AI, analytics, and settlement infrastructure that connects insurers, veterinary clinics, and pet owners.
In this briefing: 1. Briefing 01: Owning the Model, Owning the Data · 2. Briefing 02: Answers That Cite Their Sources · 3. Briefing 03: Giving the Exam Room Back · 4. Briefing 04: From Exam Room to Settled Claim · 5. Briefing 05: Predicting the Visit That Hasn’t Happened Yet · 6. Briefing 06: Engineered for Trust
Part 1: Briefing 01: Owning the Model, Owning the Data
Nearly every veterinary software vendor now claims to have AI. Almost all of them mean the same thing: a thin wrapper around a public model endpoint, with your clinical notes, your claim data, and your clients’ records leaving your control the moment a user presses Enter. Quill Direct Care made a deliberately harder engineering choice — we run the entire inference path inside our own cloud tenant, on infrastructure we provision, on networks that have no public ingress. This article explains exactly how that works, and why it is the difference between a demo and a system an insurer can underwrite against.

1. The problem with “we added AI”
The economics of the current AI market push vendors toward the fastest possible integration: sign up for a hosted API key, forward the user’s prompt, render the response. It ships in a week. It also means that a clinic’s differential diagnosis discussion, a pet owner’s financial hardship disclosure, and an insurer’s adjudication rationale all transit a third party’s multi-tenant infrastructure under whatever terms that third party publishes this quarter.
For a consumer productivity app that is an acceptable trade. For a platform that sits between a licensed veterinary practice, a regulated insurance carrier, and a pet owner’s payment instrument, it is not. Three constraints make it untenable:
-
Data residency and control. Carriers conducting vendor due diligence ask where the data physically lives, who can read it, and what contractual right the vendor has to use it. “Our AI provider says they don’t train on it” is a weaker answer than “the inference endpoint is a private endpoint on our virtual network with no route to the public internet.”
-
Auditability. When an AI-assisted recommendation influences a clinical or coverage decision, someone will eventually need to reconstruct what the model saw. That requires owning the prompt, the retrieved context, the model version, and the trace identifier — not just the answer.
-
Continuity. A pricing change, a deprecation notice, or a rate limit at a third-party API becomes an outage in a system a clinic depends on to check in patients. Owning the deployment converts a vendor risk into a capacity-planning exercise.
The architectural thesis
Model weights are a commodity and will keep getting cheaper. What is not a commodity is a production-grade, network-isolated, identity-governed, fully instrumented inference path wired into a real transactional system of record. That is what Quill built, and it is what takes eighteen months rather than eighteen days to replicate.
2. The Quill AI stack, end to end
Every AI feature in the Quill portal — the clinical scribe, the vet chat, symptom triage, coverage validation, fraud explanation, population health analysis — resolves through a single governed path. There is exactly one way for a request to reach a model, and it is instrumented.
Browser (React 18 / Vite / TypeScript)
| authenticated session, JWT bearer
v
+-----------------------------------------------------------+
| Application API (Azure Container Apps, VNet-integrated) |
| - JWKS validation against Microsoft Entra ID |
| - role + tenant authorization |
| - request tracing, cost attribution |
+-----------------------------------------------------------+
| |
| retrieval | inference
v v
+------------------------+ +------------------------------+
| Azure AI Search | | Azure OpenAI (tenant-owned) |
| vet-clinical-v1 | | gpt-5 / gpt-5-mini |
| hybrid: BM25 + vector | | text-embedding-3-small |
| PRIVATE ENDPOINT ONLY | | whisper |
+------------------------+ +------------------------------+
^ ^
| |
+-----------------------------------------------------------+
| Azure Key Vault (RBAC, purge protection, Deny-by-default)|
| Managed Identity | NAT egress pinned to a single IP |
+-----------------------------------------------------------+
Figure 1 — Request path for every AI feature in the Quill portal. No component reaches a model directly.
2.1 Tenant-owned model deployments
Quill provisions its own Azure OpenAI resource inside its own subscription and deploys named models into it — currently a reasoning-tier model, a cost-optimized model for high-volume calls, an embedding model, and a speech-to-text model. These are Quill’s deployments, addressed by Quill’s endpoint, secured by Quill’s keys and managed identity. Microsoft’s published commitment for this service is unambiguous: customer prompts, completions, and embeddings are not used to train, retrain, or improve any foundation model, and are not shared with other customers.
That commitment matters, but Quill does not rely on it alone. The endpoint is reachable only from the platform’s virtual network, authenticated by workload identity rather than a shared secret in application code, and every call is attributed to a request trace before it is issued.
2.2 The gateway abstraction
A single shared client module mediates every model call in the platform. It resolves the active gateway, normalizes model identifiers across providers, and gives Quill a switch it can throw without touching feature code:
-
One code path. Chat completions, embeddings, and audio transcription all route through the same authenticated helper. There is no second, undocumented way for a feature to call a model.
-
Model portability. Feature code requests a capability tier, not a vendor SKU. The gateway maps that tier onto whatever deployment is currently configured, so upgrading the reasoning model across thirteen backend functions is an environment-variable change, not a refactor.
-
Deterministic failover. If the primary gateway is unconfigured or unreachable, the helper degrades to a defined secondary rather than failing silently or, worse, leaking traffic to an unmanaged endpoint.
-
Cost and latency attribution. Because all traffic funnels through one place, per-feature spend and per-feature p95 latency are measurable rather than estimated.
2.3 Grounded retrieval, not open-ended generation
Quill’s clinical AI does not answer from model memory. It answers from a retrieval index and cites what it used — a design that follows directly from the retrieval-augmented generation literature, which established that grounding generation in retrieved passages produces more specific and more factual output than a parametric model alone.
The behavioral consequence is visible in production. When the index does not contain material that supports an answer, the system refuses and says so, rather than producing a plausible paragraph. In validation testing against a near-empty index, a request for a canine vaccination schedule returned a refusal and an empty citation array — exactly the intended failure mode. A veterinary AI that confidently invents a dosing interval is not a feature; it is a liability. Article 02 in this series details the retrieval architecture.
3. Governance is the product
The security posture around the AI stack is not a compliance afterthought bolted on before a sales cycle. It is the reason the stack can be sold into regulated buyers at all.
| Control | Implementation | Why a buyer cares |
|---|---|---|
| Network isolation | Azure AI Search reachable only through a private endpoint with private DNS resolution; public network access disabled at the service. | Even a leaked API key cannot be used from the public internet. |
| Secret management | All credentials in Azure Key Vault under Entra RBAC, with purge protection and soft-delete enabled; vaults set to deny-by-default with an explicit allow list. | No credential lives in source control, in an image, or in a CI variable. |
| Identity | Container Apps pull images and read secrets via user-assigned managed identity. API requests validated against Entra ID JWKS with audience pinning. | There is no long-lived password to rotate, leak, or forget. |
| Deterministic egress | All outbound traffic exits through a NAT gateway on a single static IP, which is the only address allow-listed on the database, Key Vault, and partner webhooks. | Partner allow-lists are exact, and unexpected egress is immediately visible. |
| Data-plane authorization | PostgreSQL row-level security policies enforce tenant and role scoping in the database engine, beneath the application. | An application bug cannot become a cross-tenant data breach. |
| Supply chain | CI/CD authenticates to Azure with GitHub OIDC federated credentials; images are built into a private registry and pulled by managed identity. | No cloud credential is ever stored in the CI system. |
A note on regulatory framing
Quill states this precisely, because precision is itself a due-diligence signal: HIPAA governs human protected health information and does not, as a matter of law, apply to veterinary records. Quill nevertheless engineers to HIPAA-equivalent technical safeguards and maps its control design to the SOC 2 Trust Services Criteria. The reason is practical — the carriers and enterprise clinic groups Quill integrates with run vendor security reviews written for healthcare, and the pet owner data flowing through the platform (identity, payment instruments, financial hardship, household detail) deserves that standard regardless of what the statute compels.
4. What this unlocks for each constituency
For insurers
-
A vendor whose AI data flow can be diagrammed, evidenced, and contractually bounded — which shortens security review from a blocker to a checklist.
-
Model-assisted adjudication support where every recommendation carries a trace identifier, the retrieved sources, and the prompt version that produced it.
-
No exposure to a downstream AI provider’s terms changing mid-contract, because the deployment is Quill’s.
For clinics
-
Clinical notes dictated into the scribe are transcribed by a model deployment inside the platform boundary, not uploaded to a consumer transcription service.
-
Answers cite sources, so a veterinarian can verify in seconds rather than trusting a paragraph.
-
Performance and availability are Quill’s operational responsibility, with capacity provisioned against real clinic traffic patterns.
For pet owners
-
Conversations about a pet’s health and a household’s ability to pay for it are handled as sensitive records, not as training data.
-
Guidance is grounded and cited, and the system says “I don’t know” rather than guessing about an animal’s health.
5. Why this is defensible
Competitors can license the same models tomorrow — that was always true and it is not where the advantage sits. The advantage sits in everything wrapped around the model call: private networking with working private DNS resolution, a hardened secret store, deterministic egress that partners have already allow-listed, workload identity across CI and runtime, a retrieval index with an enforced schema contract, citation guardrails validated against real refusal cases, and thirteen production features already migrated onto the shared gateway.
Each of those is individually unremarkable. Assembled, operated, and evidenced together — in a domain where the buyer is an insurance carrier’s security team — they constitute a barrier that a feature-level competitor cannot clear by adding a chat box.
Evidence base
The following peer-reviewed findings were verified against primary sources; full citations appear in the References section.
Ungrounded models fabricate citations at material rates. Across 1,800 references generated from 120 clinical anatomy questions, measured reference-hallucination rates were 23.2% for one frontier model and 45.8% and 47.5% for two others — the failure mode a citation-required architecture is designed to eliminate. (Ülkir & Paslı, Clinical Anatomy, 2026)
Part 2: Briefing 02: Answers That Cite Their Sources
The single most dangerous property of a general-purpose language model in a clinical setting is that it is equally fluent when it is right and when it is wrong. Quill’s clinical AI is built so that fluency is decoupled from authority: an answer is only rendered when retrieved source material supports it, every claim carries a numbered citation back to a document, and the absence of grounding produces a refusal rather than a guess.

1. Why grounding, not fine-tuning
There are two ways to give a model domain knowledge. You can bake it into the weights through fine-tuning, or you can retrieve it at query time and place it in context. Quill chose retrieval, and the reasoning is operational as much as technical.
| Dimension | Fine-tuned model | Retrieval-augmented (Quill’s approach) |
|---|---|---|
| Updating knowledge | Retrain and redeploy; days to weeks. | Re-run the ingestion job; minutes. |
| Attribution | None. The model cannot tell you which document taught it something. | Every sentence maps to a retrieved chunk with a document identifier and URL. |
| Correcting an error | The bad fact is diffused across the weights. | Fix or remove the source document; the error disappears on the next query. |
| Access control | Knowledge is baked in for everyone who can call the model. | Retrieval is filtered per request, so scoping is enforceable. |
| Failure mode | Confident fabrication. | Explicit refusal with an empty citation array. |
The retrieval-augmented generation approach was introduced precisely because grounding generation in retrieved passages produces more specific and more factual output than a parametric model relying on memorized weights. Quill’s implementation adds a hard constraint on top of it: no citation, no answer.
2. How a question becomes a cited answer
Question
|
+--> [1] Embed the query -> text-embedding-3-small (1536-dim)
|
+--> [2] Hybrid retrieval -> Azure AI Search vet-clinical-v1
| a. vector similarity (semantic meaning)
| b. BM25 keyword match (exact drug names, dosages, codes)
| c. reciprocal rank fusion of both result sets
|
+--> [3] Filter + rerank -> species, topic, recency, permission
|
+--> [4] Assemble numbered context -> [1] .. [N] with doc_title + url
|
+--> [5] Generate under constraint -> gpt-5-mini, citation-required prompt
|
+--> [6] Validate -> every claim maps to a source, or refuse
|
v
{ answer, citations[], prompt_version, retrieval_trace_id }
Figure 1 — The retrieval pipeline behind every clinical answer.
2.1 Why hybrid retrieval and not pure vector search
Pure semantic search is excellent at concepts and unreliable at tokens. Ask it for content about “kidney problems in older cats” and it performs beautifully. Ask it for a specific compound, a product name, a dosage string, or a claim code and dense embeddings will happily return something adjacent and wrong — because in embedding space, two drug names in the same class sit very close together.
Veterinary medicine is full of exactly those tokens. Quill therefore runs vector and keyword retrieval in parallel and fuses the ranked lists with reciprocal rank fusion, so conceptual recall and literal precision reinforce each other rather than trading off. A semantic reranking pass then reorders the fused set by relevance to the actual question rather than by raw score.
2.2 The schema contract
The retrieval index enforces a strict field contract — document identifier, source, document title, section, species, topics, canonical URL, publication date, content, and content vector. The ingestion pipeline reads the live index schema at startup and drops any field the index does not define before it writes.
That guard exists because of a real production failure. An early ingestion run emitted a title field against an index that defines doc_title, and the service correctly rejected the batch. Rather than patching the one field name, Quill made the pipeline schema-aware so the entire class of mismatch cannot recur. This is the difference between fixing a bug and closing a defect category — and it is the kind of hardening that only comes from operating a pipeline rather than demoing one.
2.3 The citation guardrail
The generation step runs under a versioned system prompt that forbids unsourced assertions. The response contract is machine-checkable:
-
answer — prose in which every substantive claim carries a bracketed numeric citation.
-
citations[] — the resolved source list: document title, section, and canonical URL for each number used.
-
prompt_version — the exact prompt revision that produced the answer, so behavior changes are attributable.
-
retrieval_trace_id — a unique identifier linking the answer to the precise retrieval result set, for audit and evaluation replay.
The refusal is the feature
During validation of the live endpoint, a request for a puppy vaccination schedule was issued against an index containing only a seed document. The system returned a refusal, an empty citations array, and a clarifying question — it did not produce a schedule from model memory.
That is the behavior clinical users need and the behavior a carrier’s risk team wants to see demonstrated. An AI that never says “I don’t know” is an AI that cannot be trusted when it says anything else.
3. Curation is a governance decision, not a scraping job
Retrieval quality is bounded by corpus quality, so Quill treats corpus admission as a controlled process. Content enters the index through a versioned manifest that records provenance, license status, publication date, and review owner for every document. Material without a clean license or a clear provenance chain does not get ingested, regardless of how useful it would be.
Ingestion itself runs as a scheduled, event-driven containerized job that scales to zero between runs — it costs nothing when idle, and each execution emits a log line with the document count written, so a silent partial ingest is detectable. Refresh is a job run, not a deployment.
4. Where grounded answers show up in the product
| Surface | User | What grounding changes |
|---|---|---|
| AI Vet Chat | Pet owner | Guidance is tied to the specific pet’s species, age, breed, and history — and cites the material behind each recommendation instead of producing generic web-tier advice. |
| Symptom triage | Pet owner | Urgency assessment is conservative by construction: when signals are ambiguous, the system escalates to “be seen” rather than reassuring. |
| Diagnostic Encyclopedia | Veterinarian | Presents multiple valid diagnostic approaches with the reasoning behind each, explicitly framed as decision support rather than protocol. |
| Coverage validation | Clinic staff, insurer | Coverage determinations reference the specific policy language relied on, so a denial can be explained to an owner in plain terms. |
| Fraud explanation | Insurer | A flag is accompanied by the pattern and the evidence that produced it, which makes the signal reviewable rather than oracular. |
| Population health | Medical director | Cohort findings link back to the underlying records, so a practice can act on a trend instead of debating whether it is real. |
5. Evaluation as a standing discipline
A grounded system still degrades — an index drifts, a prompt revision regresses, a model upgrade changes refusal calibration. Quill instruments for that:
-
Retrieval quality. Held-out question sets measure whether the correct source appears in the top-k results, tracked per content domain.
-
Citation fidelity. Answers are checked so that every cited number resolves to a retrieved document, and every substantive claim carries one.
-
Refusal calibration. A deliberate set of unanswerable questions verifies the system still declines — over-refusal and under-refusal are both regressions.
-
Latency and cost. Per-feature p95 latency and per-answer token spend are tracked at the gateway, so a quality improvement that quietly triples cost is visible immediately.
-
Replay. Because every answer carries a trace identifier and a prompt version, any historical answer can be reconstructed and re-scored against a new configuration.
6. Why this matters for care
Veterinary practice runs on compressed decision time. A clinician holding a differential at 4:40 p.m. does not need a fluent essay — they need the two or three relevant considerations and a source they can check in fifteen seconds. A pet owner at 11 p.m. deciding whether to drive to an emergency clinic does not need reassurance — they need a conservative, explainable read on urgency.
Citations serve both. They convert the AI from an oracle that must be trusted into an instrument that can be verified, and verification is what makes it safe to use the tool at speed. That is the entire design goal: not to replace veterinary judgment, but to make the evidence supporting a judgment reachable in seconds instead of minutes.
Evidence base
The following peer-reviewed findings were verified against primary sources; full citations appear in the References section.
Ungrounded models fabricate citations at material rates. Across 1,800 references generated from 120 clinical anatomy questions, measured reference-hallucination rates were 23.2% for one frontier model and 45.8% and 47.5% for two others — the failure mode a citation-required architecture is designed to eliminate. (Ülkir & Paslı, Clinical Anatomy, 2026)
Retrieval grounding measurably reduces hallucination. Adding a retrieval-augmented generation framework to four large language models across four medical question-answering benchmarks increased accuracy by 9.8–16.3% (mean 13%) and reduced hallucination rates by 11.8–18% (mean 15.1%). (Wang et al., Journal of Medical Internet Research, 2025)
Part 3: Briefing 03: Giving the Exam Room Back
The most expensive resource in a veterinary practice is a licensed clinician’s attention, and the industry spends an extraordinary share of it on typing. Quill’s AI Scribe listens to the visit, separates who said what, and produces a structured, editable clinical note that flows directly into the chart, the treatment plan, the estimate, and — when relevant — the insurance claim. Nothing is auto-filed. Everything arrives pre-drafted.

1. The documentation tax
Every veterinary practice pays a documentation tax, and it compounds in three places. Notes written during the visit fracture the client conversation. Notes written after the visit extend the working day past its scheduled end. Notes deferred to the end of the week arrive thinner, later, and less useful to whoever reads the chart next — including the insurer who eventually needs the medical record to adjudicate a claim.
The downstream cost is not just clinician time. Incomplete documentation is one of the most common reasons a legitimate claim stalls in review. A note that omits the presenting complaint, the date of onset, or the diagnostic rationale forces a carrier to request records, which adds days to reimbursement and lands the delay on the pet owner. Documentation quality and financial throughput are the same problem viewed from two ends.
2. How the scribe works
Exam room audio (browser capture, explicit consent, visible recording state)
|
v
[1] Transcription -> Whisper deployment inside the Quill tenant
| timestamped segments, robust to noise and accents
v
[2] Diarization -> speaker separation: clinician / technician / owner
|
v
[3] Trigger matching -> phrase library + built-in triggers detect
| vitals, meds, vaccines, findings, follow-ups
v
[4] Structuring -> SOAP: Subjective / Objective / Assessment / Plan
| coded to the practice's preset template
v
[5] Clinician review -> side-by-side draft + transcript, inline edit
| NOTHING enters the chart unsigned
v
Chart -> Treatment plan -> Estimate -> Claim packet
Figure 1 — The ambient documentation pipeline. Every stage is reviewable and reversible.
2.1 Transcription that survives the exam room
Exam rooms are acoustically hostile: barking, clipper noise, a scale beeping, two people talking over each other, and a vocabulary dense with drug names and anatomical terms. Quill runs a large-scale weakly-supervised speech recognition model whose defining property is robustness to exactly these conditions — accents, background noise, and technical vocabulary — deployed inside Quill’s own cloud tenant rather than sent to a consumer transcription service.
That deployment choice is not incidental. Exam room audio contains a client’s financial circumstances, a clinician’s uncertainty, and occasionally a difficult end-of-life conversation. It belongs inside the platform boundary, on infrastructure where the operator has published and contractually bound commitments that the content will not be used to train foundation models.
2.2 Diarization: who said what
A transcript without speaker attribution is nearly useless for clinical documentation, because the Subjective section is by definition what the owner reported and the Objective section is what the clinician observed. Quill separates speakers so the structuring stage can route content to the right part of the note instead of producing an undifferentiated wall of text that a clinician then has to reorganize by hand — which would defeat the purpose entirely.
2.3 Triggers and phrase libraries
Practices do not document identically, and a scribe that imposes one house style gets abandoned in a month. Quill ships a library of built-in triggers and presets covering common visit types, and lets a practice extend both:
-
Built-in triggers recognize clinically salient utterances — a stated weight, a temperature, a vaccine administered, a medication and dose, a recheck interval — and place them into the correct structured field rather than leaving them buried in prose.
-
Presets encode a visit archetype (wellness, dental, dermatology, orthopedic recheck, euthanasia consult) so the resulting note matches the shape the practice already uses.
-
Custom phrase libraries let a practice add its own vocabulary, abbreviations, and standard language, which is what turns a generic tool into that practice’s tool.
2.4 The clinician is always the author
Design principle: draft, never file
The scribe produces a draft. A licensed clinician reads it, edits it, and signs it. There is no configuration in which an AI-generated note enters the legal medical record unreviewed.
The transcript remains available alongside the draft, so a clinician can verify any line against what was actually said rather than trusting the summary. Review is fast because it is verification, not authorship.
3. What the note connects to
A scribe that produces a good note and stops has solved a typing problem. Quill’s scribe sits inside a transactional system, so the structured note becomes an input to everything downstream of the visit.
| Downstream artifact | What the structured note contributes | Who benefits |
|---|---|---|
| Patient chart | Complete SOAP entry with vitals, findings, and assessment already coded to fields rather than free text. | The next clinician to open the chart |
| Treatment plan | Plan items lifted from the Plan section as discrete, priceable line items with clinical rationale attached. | Clinician and owner |
| Cost estimate | Line items priced against the practice fee schedule before the client leaves the room. | Owner — no surprise invoice |
| Claim packet | Presenting complaint, onset date, diagnostic findings, and rationale — the exact fields carriers request when they hold a claim. | Owner and insurer |
| Follow-up and recall | Recheck intervals detected in the conversation become scheduled outreach with consent applied. | Practice revenue and patient outcomes |
| Population analytics | Coded findings aggregate into cohort-level trends rather than being locked in prose. | Medical director |
This is the compounding effect that a standalone scribe product cannot produce. When documentation, plan, estimate, and claim are separate systems, the clinician’s note has to be re-keyed three times by three people, and each re-keying is an opportunity for the claim to lose the detail the carrier needs. Inside Quill they are one object viewed from four angles.
4. Quality, safety, and the audit trail
-
Consent and visibility. Recording state is explicit in the interface. Capture is a deliberate act, not an ambient default.
-
Retention control. Audio and transcript retention are policy-configurable per practice, with the underlying records governed by row-level security in the database.
-
Full attribution. Every generated draft carries the model deployment, prompt version, and trace identifier that produced it, so a note can be reconstructed for audit years later.
-
Edit history. The delta between generated draft and signed note is preserved — which is both a medico-legal record and the most valuable quality signal the system produces.
-
Closed-loop improvement. Systematic clinician corrections identify where the structuring is weak, which drives prompt and trigger revisions rather than guesswork.
5. The operational case
The argument for ambient documentation is usually made as time saved per note. That is the smallest part of the value. The larger effects are structural:
-
Attention returns to the client. A clinician who is not typing is looking at the animal and the person who brought it, which is where diagnostic signal and client trust both come from.
-
Notes get written at the point of care. Same-day documentation is more accurate than end-of-week documentation, and accuracy is what determines whether a claim clears on first submission.
-
Claims clear faster. Complete presenting complaint, onset, and rationale reduce the records requests that stall reimbursement — which the owner experiences as getting their money back sooner.
-
Charts become data. Structured capture means a practice can answer questions about its own population instead of guessing, which is the precondition for the analytics described in Article 05.
-
Retention improves. Administrative load is a documented driver of attrition in clinical professions. Removing it is a staffing intervention with a measurable cost basis, not a soft benefit.
None of this replaces clinical judgment, and Quill does not market it that way. The scribe removes transcription labor so that judgment has more room to operate. That is a narrower claim than the market usually makes, and it is one the product actually delivers.
Evidence base
The following peer-reviewed findings were verified against primary sources; full citations appear in the References section.
Clinician distress is measured, not anecdotal. In a nationwide cross-sectional study of 2,208 US veterinary-profession respondents, 888 (41%) reported serious psychological distress and 381 (17.3%) had considered suicide in the previous 12 months. (Scoresby et al., Frontiers in Veterinary Science, 2023)
Documentation time is the intervention target. A five-site study of 8,581 clinicians treated electronic-record documentation burden as the measurable baseline problem, quantifying total record time and documentation time per eight scheduled patient hours. (Rotenstein et al., JAMA, 2026)
Ambient AI scribes return measurable clinical time. In a difference-in-differences analysis across five US academic health systems, scribe adoption was associated with 13.4 fewer minutes of total record time (95% CI 9.1–17.7) and 16.0 fewer minutes of documentation time (95% CI 13.7–18.3) per eight scheduled patient hours, plus 0.49 additional weekly visits (95% CI 0.17–0.81). (Rotenstein et al., JAMA, 2026)
Part 4: Briefing 04: From Exam Room to Settled Claim
Pet insurance has a settlement problem, and it is not an actuarial one. The claim is submitted after the fact, by the owner, from a receipt, into a carrier process the clinic cannot see — and the owner floats the entire cost until it resolves. Quill collapses that sequence: coverage is validated before treatment, the claim is assembled from the clinical record rather than reconstructed from paper, and the money moves along a state machine and a double-entry ledger that both sides can inspect.

1. The structural defect in the current flow
CONVENTIONAL
Visit -> Owner pays 100% -> Owner finds receipt -> Owner files claim
-> Carrier requests records -> Clinic mails records -> Adjudication
-> Reimbursement [days to weeks of float]
QUILL DIRECT CARE
Pre-visit coverage validation
-> Visit + ambient documentation
-> Estimate priced against the fee schedule, owner sees split up front
-> Claim assembled from the clinical record, submitted at checkout
-> Adjudication with complete evidence on first pass
-> Owner pays their share only [float removed or absorbed]
Figure 1 — Conventional reimbursement versus the Quill path.
The pet insurance market is growing quickly — NAPHIA’s State of the Industry reporting, drawn from members representing 99% of North American policies in effect, tracks a steadily expanding insured-pet population and sustained double-digit annual growth in gross written premium. Growth of that speed makes the settlement defect more expensive every year, not less: more policies flowing through a reimbursement model that was designed around mailed receipts.
The defect has three distinct costs. The owner carries the full price of care at the moment of highest emotional stress. The clinic absorbs the collections risk and the records-request labor. The carrier adjudicates from a reconstructed narrative rather than the primary clinical record, which is why records requests exist at all.
2. The claim as a state machine
Quill does not treat claim status as a text field that services update opportunistically. It is a formal state machine with defined states and explicitly permitted transitions, implemented in code and covered by tests. An illegal transition is rejected, not logged and tolerated.
| Stage | Representative states | What is guaranteed |
|---|---|---|
| Pre-service | draft, coverage_validated, awaiting_visit | Eligibility and expected coverage are established before treatment begins, so the estimate the owner sees is real. |
| Submission | submitted, processing | The claim carries the clinical record, the itemized invoice, and the evidence attachments as one packet. |
| Adjudication | processing, information_requested, approved, denied | Every determination is attributable to a rule or reviewer, with the rationale retained. |
| Settlement | approved, paid, settled | Ledger entries post in balanced pairs; a payment cannot exist without its counterpart. |
| Exception | information_requested, disputed, cancelled | Exceptions route to a queue with an owner and an age, rather than disappearing into a status. |
Why a state machine is a commercial asset
Because transitions are explicit, every claim has a complete, ordered, timestamped history. That single property produces auditability for the carrier, real status visibility for the owner, cycle-time analytics for the practice, and a testable invariant for the engineering team.
It is also what makes exception handling tractable: an exception is a claim that has been in a state too long, which is a query — not an investigation.
3. Three payment paths, one transparent engine
Quill’s fee engine calculates every transaction through one tiered schedule, and it is deliberately legible. The fee is computed on the amount Quill is actually facilitating, tiered by transaction size, bounded by configured minimum and maximum, and — critically — disclosed to every party before anyone commits.
| Path | Mechanics | Who carries the float | Clinic receives |
|---|---|---|---|
| Direct Pay | The insurer-covered portion is paid to the clinic directly; the owner pays only their responsibility. | Insurer | Full service total |
| Quill Advance | Quill fronts the insurer-covered portion at the point of care and recovers it from the carrier. | Quill | Full service total |
| Owner Self-Pay | Uninsured or non-covered service, facilitated through the same rails. | Owner | Service total less fee |
Two properties of the engine matter more than the rates. First, the clinic receives the full service total on both insured paths — the practice is not financing the insurance industry’s processing time. Second, on facilitated paths the platform fee is split transparently rather than embedded in a spread, and the breakdown is rendered to the owner in the checkout flow. Fee schedules are configurable per practice and versioned, so a rate change is an auditable event rather than a silent repricing.
Worked example
A $1,200 procedure with $900 of insurer-covered benefit under Direct Pay: the owner’s responsibility is the $300 balance plus a disclosed share of the platform fee calculated on the $900 facilitated amount; the clinic receives the full $1,200; the carrier settles the $900 through the ledger. Every number in that sentence is visible to all three parties in the interface before the transaction is authorized.
4. The ledger
Money movement in Quill is recorded as double-entry accounting, not as status flags on a payments table. Every advance, recovery, fee, refund, and settlement posts as balanced entries against defined accounts.
-
Reconciliation is arithmetic. Carrier remittances reconcile against expected recovery amounts automatically; the exception queue holds only genuine mismatches.
-
Recovery is tracked as an asset. Quill Advance creates a receivable with an explicit aging profile, which is what makes the advance product underwritable rather than a hope.
-
Financial state is derived, never asserted. Balances are computed from entries. There is no field a bug can set to the wrong number.
-
Audit reconstruction is complete. Any balance at any historical moment can be rebuilt from the entry log, which is the baseline expectation of a SOC 2 examination.
-
Tenant isolation is enforced below the application. PostgreSQL row-level security policies scope financial records in the database engine, so an application-layer defect cannot expose another party’s ledger.
5. Meeting carriers on their own rails
Quill does not require a carrier to adopt a new integration standard as the price of participation — a requirement that has killed more insurtech pilots than any technical failure. The exchange layer supports the protocols carriers actually run:
| Protocol | Typical carrier posture | Quill support |
|---|---|---|
| REST / JSON | Modern digital-native carriers and MGAs. | Native, with webhook callbacks for status transitions. |
| Webhooks | Event-driven status and remittance notification. | Signed, retried, and observable — with delivery metrics and a replay path. |
| EDI X12 (837 / 835 / 270 / 271) | Established carriers with mature clearinghouse infrastructure. | Claim submission, remittance advice, and eligibility inquiry mapped to the standard sets. |
| HL7 / FHIR R4 | Carriers and partners aligned to healthcare interoperability. | Resource-level mapping for clinical evidence exchange. |
| SFTP batch | Carriers operating scheduled batch reconciliation. | Scheduled drops and pickups with manifest validation. |
| ETL buckets | Analytics and bulk data-warehouse feeds. | Structured extracts on a defined cadence. |
Each pathway resolves into the same internal claim object and the same state machine. Protocol is an edge concern; the core is uniform. That is what allows Quill to onboard a carrier on the carrier’s timeline rather than on a platform migration timeline.
6. Integrity and exception handling
Faster settlement without integrity controls is just faster leakage. Quill pairs acceleration with detection:
-
Pre-service validation catches coverage problems before treatment, which eliminates the worst outcome in the system — a denial discovered after the owner has already been told they were covered.
-
Evidence at submission. Receipts and records are attached and OCR-extracted at the point of claim assembly, and the extracted values are cross-checked against the itemized invoice.
-
Pattern-level fraud signals operate across claims rather than within one, and each flag is delivered with the pattern and evidence that generated it so a human reviewer can adjudicate the flag itself.
-
Aged exception queues hold anything that stalls, with an owner and a clock, so no claim quietly ages out of anyone’s attention.
-
Immutable audit trail. Every state transition, review decision, and ledger posting is recorded with actor and timestamp, and the underlying database is deployed with zone-redundant high availability and point-in-time restore.
7. Why this is the hard part to copy
A competitor can build a claim submission form in a sprint. What takes years is the composite: a tested state machine that carriers will accept as authoritative, a configurable fee engine with disclosed economics, a double-entry ledger that survives an audit, six live exchange protocols so no carrier has to change to participate, and pre-service coverage validation wired into the clinical workflow that generates the claim in the first place.
Each layer is only valuable in the presence of the others. Coverage validation without clinical integration is a lookup tool. A ledger without settlement volume is a spreadsheet. Multi-protocol exchange without a normalized claim object is six brittle integrations. Assembled, they are a payments and adjudication network for a market growing at double-digit annual rates — and the participants on both sides are already on it.
Evidence base
The following peer-reviewed findings were verified against primary sources; full citations appear in the References section.
Cost governs the care a pet actually receives. Of 2,296 pet guardians surveyed, 1,924 (83.9%) agreed that the expense of veterinary care affects the level of healthcare their pets receive, while only about 36% considered pet insurance essential. (Awawdeh, Waran & Forrest, Animal Welfare, 2025)
Part 5: Briefing 05: Predicting the Visit That Hasn’t Happened Yet
Veterinary medicine and pet insurance are both structurally reactive: care is delivered when an animal presents, and claims are paid when care is delivered. Quill sits on the one dataset that can change that — longitudinal clinical records, claim histories, and owner engagement in a single system — and turns it into a per-pet risk score, a per-practice population view, and a per-cohort actuarial signal.

1. Why the data has never existed in one place
The information needed to anticipate a pet’s next health event has always been fragmented across parties that do not exchange it. The clinic holds the clinical record. The carrier holds the claim history. The owner holds adherence, diet, environment, and everything that happens between visits. No participant can see the whole animal.
Quill is a transactional system of record for all three constituencies simultaneously. Because clinical documentation, claim submission, scheduling, and owner engagement all flow through the same platform, the longitudinal picture assembles as a byproduct of normal use rather than as a data-warehouse project nobody funds.
This is a consequence of the transaction model
Quill’s analytics are not a reporting module bolted onto a records system. They are a second-order effect of being the transaction rail. Every visit documented, every claim submitted, and every reminder acknowledged deepens the same longitudinal record — which means the predictive layer improves as a function of platform usage, not engineering spend.
2. The per-pet risk engine
Quill scores each pet on a 0–100 composite drawn from independently computed dimensions. The score is explicitly decomposable — a number without its contributing factors is not clinically actionable, and would not be accepted by a veterinarian.
| Dimension | Signals evaluated | Clinical significance |
|---|---|---|
| Age and breed | Life stage against species norms; breed-linked predispositions; healthy-weight range for breed and sex. | Establishes the baseline hazard profile before any individual history is considered. |
| Clinical history | Prior diagnoses, chronic conditions, procedures, and recorded findings over time. | Past disease is the strongest available predictor of future disease. |
| Claims pattern | Frequency, severity, and clustering of claims relative to peers. | Escalating utilization frequently precedes a chronic diagnosis. |
| Preventive compliance | Vaccination currency, wellness visit cadence, recheck adherence, medication refills. | Adherence gaps are the most directly modifiable risk factor in the model. |
| Lifestyle and environment | Weight trajectory against breed range, diet, activity, household context, temperament. | Weight trend in particular is an early, measurable, and reversible signal. |
2.1 Explainability by construction
Every score renders with its dimensional breakdown and a set of specific recommendations attached to the dimensions that drove it. A clinician sees not “this pet is 72” but which factors produced 72 and which of them can be changed. That design constraint — no opaque scores — is what determines whether the tool is used or ignored after week three.
2.2 From score to action
-
Detect. The engine surfaces a rising trajectory — a weight trend crossing the breed range, a lapsed preventive interval, a claim pattern consistent with an emerging chronic condition.
-
Recommend. Concrete next steps are generated and attached to the specific dimension that triggered them, so the recommendation carries its own justification.
-
Route. Recommendations reach the right actor: the pet owner’s timeline, the clinic’s care management queue, or a consent-filtered outreach campaign.
-
Schedule. Accepted recommendations convert into appointments through the scheduling command center, where AI-assisted slotting reduces the friction between intent and a booked visit.
-
Close the loop. Outcomes feed the treatment-effectiveness view, so the practice learns which interventions actually changed the trajectory rather than assuming they did.
The last step is the one most predictive products omit. A prediction that is never checked against an outcome is a marketing asset. Quill measures whether the recommended intervention was delivered, and whether the risk trajectory subsequently changed.
3. Population health at the practice level
Aggregating individual risk across a practice’s patient base produces a view no practice has historically been able to construct without a data analyst:
-
Cohort risk distribution — how many patients sit in each risk band, and how that distribution is moving quarter over quarter.
-
Preventive care gaps — the specific list of patients overdue for a specific intervention, ranked by risk contribution rather than alphabetically.
-
Benchmarking — practice performance against comparison cohorts on compliance, chronic disease prevalence, and outcome measures.
-
Chronic disease burden — prevalence trends that inform staffing, equipment, and referral relationships.
-
Revenue and capacity implications — the forecast visit demand implied by the population’s risk profile, which turns clinical insight into a scheduling and hiring plan.
Reports are generated with AI assistance and exported as clinical documents for practice leadership, with history retained so a medical director can compare this quarter to last rather than re-deriving the baseline each time.
4. What this means for insurers
Carriers price risk on enrollment data and pay claims on submitted evidence. Between those two moments they have historically been blind. A platform that observes the clinical trajectory continuously changes several things at once:
| Capability | Mechanism | Carrier impact |
|---|---|---|
| Earlier intervention | Preventive gaps identified and closed before an acute presentation. | Lower severity per claim; fewer emergency presentations. |
| Cohort-level loss signals | Risk distribution shifts observable across a book, not just at renewal. | Reserving and pricing informed by leading indicators rather than trailing loss runs. |
| Adherence measurement | Whether recommended care was actually delivered, measured rather than surveyed. | Wellness product design that can be validated instead of assumed. |
| Fraud and anomaly context | Utilization patterns evaluated against a longitudinal baseline for that animal. | Fewer false positives; flags that survive review. |
| Product differentiation | Behavior-linked and preventive-linked benefit structures become measurable. | New products in a market with sustained double-digit annual growth. |
None of this requires the carrier to build the data platform. It requires the carrier to participate in one that already sits at the point of care.
5. Guardrails
Predictive systems in health domains fail in predictable ways, and Quill is engineered against each:
-
Decision support, never determination. Risk scores inform clinical and coverage judgment. No score automatically denies coverage or dictates treatment.
-
Transparency to the clinician. Every score exposes its inputs. A clinician who disagrees can see exactly which dimension is driving the number and override it.
-
Bias surveillance. Breed and age are legitimate clinical variables and also plausible vectors for unfair outcomes. Model behavior is monitored across breed, age, and geography for disparities that are not clinically justified.
-
Data isolation. Practice and owner data are scoped by row-level security in the database engine. Cross-practice benchmarking uses aggregates, not readable records.
-
Owner agency. Pet owners see their own pet’s risk profile and the reasoning, rather than being scored invisibly by a system that only their insurer can read.
6. The outcome that matters
Pets do not benefit from a risk score. They benefit from the visit that happens three months earlier than it otherwise would have — the weight trend caught before it became endocrine disease, the dental disease addressed before extraction, the lapsed preventive interval closed before an infection.
Every component in this article exists to make that earlier visit happen: detect the trajectory, explain it to a clinician who can act on it, reach the owner through a channel they consented to, make booking trivial, and then measure whether it worked. That is the whole thesis — and it is only executable from a platform that is simultaneously the clinical record, the claim rail, and the owner’s front door.
Evidence base
The following peer-reviewed findings were verified against primary sources; full citations appear in the References section.
Early detection depends on structured wellness data. Machine-learning classification of wellness visits was developed precisely because early disease detection in companion animals depends on identifying subclinical abnormalities recorded during routine preventive care — which requires those visits to be reliably identifiable in the record. (Szlosek et al., Frontiers in Veterinary Science, 2024)
Cost governs the care a pet actually receives. Of 2,296 pet guardians surveyed, 1,924 (83.9%) agreed that the expense of veterinary care affects the level of healthcare their pets receive, while only about 36% considered pet insurance essential. (Awawdeh, Waran & Forrest, Animal Welfare, 2025)
Part 6: Briefing 06: Engineered for Trust
In regulated distribution, the security review is the sales cycle. A platform that sits between a licensed practice, a regulated carrier, and a consumer’s payment instrument does not get evaluated on its interface — it gets evaluated on whether a vendor risk team can approve it. Quill Direct Care was architected with that review as a design input, which is why the controls in this article are structural rather than compensating.

1. The threat model
Quill’s platform concentrates four categories of sensitive data that are rarely found together, and the architecture is organized around that concentration:
| Data class | Examples | Primary risk |
|---|---|---|
| Clinical records | Diagnoses, notes, imaging, lab results, exam room audio. | Disclosure, tampering, loss of medico-legal integrity. |
| Financial instruments | Payment methods, advances, receivables, ledger entries. | Fraud, unauthorized movement, reconciliation failure. |
| Personal data | Owner identity, contact details, household context, financial hardship. | Disclosure, unlawful secondary use. |
| Carrier data | Policy terms, adjudication rules, remittance detail. | Competitive exposure, contractual breach. |
2. Network architecture: no public data plane
The controlling decision in Quill’s infrastructure is that the data plane has no public presence. Application services run inside a virtual network; the data services they depend on are reachable only from inside it.
Internet
|
[ HTTPS ingress ]
|
+====================== VIRTUAL NETWORK =========================+
| | |
| +----------------------v-----------------------+ |
| | Container Apps environment (VNet-integrated) | |
| | API | RAG service | ingestion jobs | |
| +---+--------+--------------+-------------------+ |
| | | | |
| private private private egress |
| endpoint endpoint endpoint | |
| | | | v |
| +----v---+ +--v-------+ +----v-----+ +-----------+ |
| | Key | | AI Search| | Postgres | | NAT | |
| | Vault | | (no pub) | | flexible | | gateway | |
| +--------+ +----------+ +----------+ +-----+-----+ |
| | |
| Bastion --> jump host (no public SSH) | |
+===============================================|=================+
v
single static egress IP
(the only address on every partner allow-list)
Figure 1 — Network topology. Only the application ingress is publicly addressable.
-
Private endpoints with private DNS. Search and the vault resolve to private addresses inside the VNet. Public network access is disabled at the service, so a stolen key is unusable from outside — a property Quill has verified operationally, not just configured.
-
Deterministic egress. All outbound traffic leaves through a NAT gateway on one static IP. That address is the sole entry on the database firewall, the vault allow-list, and partner webhook allow-lists. Unexpected egress is therefore immediately anomalous.
-
Deny-by-default vaults. Key Vault network rules default to Deny with an explicit allow list, with soft-delete and purge protection enabled so a destructive action is recoverable.
-
No public administrative access. Database and host administration run through a bastion-tunneled jump host. There is no internet-facing SSH port and no publicly reachable database endpoint.
3. Identity: credentials that do not exist cannot leak
The most common cause of cloud compromise remains a credential in the wrong place — a key in a repository, a connection string in an image, a service principal secret in a CI variable. Quill’s design goal is to minimize the number of long-lived secrets that exist at all.
| Layer | Mechanism | Secrets eliminated |
|---|---|---|
| User authentication | Microsoft Entra ID with enforced MFA; JWKS validation with audience pinning at the API. | No application-managed password store for staff identities. |
| Service-to-service | User-assigned managed identity for registry pulls and vault reads. | No registry password, no vault client secret. |
| CI/CD | GitHub OIDC federated credentials scoped to repository and branch. | No cloud credential stored in the CI system at all. |
| Application secrets | Key Vault references resolved at container start under RBAC. | No secret in source, image layers, or environment files. |
| Contractor access | Entra guest identities in a scoped group, time-bounded, least-privilege RBAC. | No shared accounts, no standing broad access. |
Least privilege, applied to people as well as services
External contractors receive guest identities placed in a purpose-scoped group with only the roles their work requires, reaching the database exclusively through a bastion tunnel and a matching database role. Access is granted per engagement and revoked with the engagement — a group membership change, not a credential rotation exercise.
4. Authorization enforced by the database
Application-layer authorization is necessary and insufficient. A single missing filter in one query handler is a cross-tenant breach in most multi-tenant systems. Quill enforces scoping one layer lower, in PostgreSQL row-level security policies evaluated by the database engine itself.
-
Policies scope rows by authenticated identity, tenant, and role for every table holding clinical, financial, or personal data.
-
Role membership is held in a dedicated roles table — never as a mutable attribute on a user or profile record — and evaluated through a security-definer function, which is the standard mitigation against privilege-escalation via self-service profile edits.
-
Explicit grants are issued per table per role, so access is affirmative rather than inherited by accident.
-
Personally identifiable fields are masked in interfaces where full values are not operationally required.
-
The result is a hard floor: an application bug becomes an error, not a disclosure.
5. Auditability
Every consequential action in the platform produces an immutable record with actor, timestamp, and before/after state — clinical note signature, claim state transition, ledger posting, role assignment, coverage determination, and AI-assisted recommendation alike.
AI actions are audited to the same standard as human ones. Each generated output carries its model deployment, prompt version, retrieval trace identifier, and the resolved source set, so an answer produced eighteen months ago can be reconstructed exactly. Auditability of automated decisions is rapidly becoming a procurement requirement rather than a differentiator, and retrofitting it is substantially harder than designing for it.
6. Resilience and operations
-
Database durability. Zone-redundant high availability with point-in-time restore on the managed PostgreSQL service.
-
Stateless compute. Application containers hold no durable state, so a revision can be replaced or rolled back without data risk.
-
Immutable, versioned deployments. Every release is a tagged image in a private registry pulled by managed identity; rollback is a revision pointer change.
-
Scale-to-zero batch. Ingestion and maintenance run as event-driven container jobs that consume nothing when idle — a cost control that is also a reduced attack surface.
-
Cost governance as a control. Budget alerts, log-ingestion caps, and periodic spend review are treated as operational controls; unmonitored spend is how surprise infrastructure appears, and surprise infrastructure is how unmonitored exposure appears.
-
Documented configuration. The full infrastructure configuration — including the exact commands used and the reasoning behind each decision — is maintained as a living operational runbook, which is what makes disaster recovery a procedure rather than an archaeology project.
7. Compliance posture, stated precisely
Quill is deliberate about what it claims, because overstatement is itself a due-diligence failure:
-
HIPAA governs human protected health information and does not, as a matter of law, extend to veterinary records. Quill engineers to HIPAA-equivalent technical safeguards because its buyers’ security programs are written that way and because the personal and financial data in the platform warrants it — not because a statute compels it for animal health data.
-
SOC 2 Trust Services Criteria provide the control framework Quill’s design is mapped to: security, availability, processing integrity, confidentiality, and privacy.
-
Payment data is handled through certified processors; Quill’s ledger records movement and reconciliation rather than storing raw instrument data.
-
AI data handling rests on a published commitment from the platform operator that prompts, completions, and embeddings are not used to train or improve foundation models, reinforced by Quill’s own network isolation so the commitment is not the only control.
8. Why this is commercially decisive
Security architecture is usually presented as a cost center. In this market it is the distribution channel. A carrier’s vendor risk team will ask where the data lives, how identity is federated, whether the data plane is publicly reachable, how automated decisions are audited, and how contractor access is scoped and revoked. A vendor who answers those questions with a diagram and evidence moves to contracting. A vendor who answers with intentions does not move at all.
Every control described here already exists in Quill’s deployed environment, was implemented and verified operationally rather than asserted in a policy document, and is recorded in a configuration runbook with the exact commands and reasoning behind each decision. That evidence base is not a byproduct of the build — it is one of the assets.
Consolidated references
1. Microsoft, “Data, privacy, and security for Azure OpenAI in Azure AI Foundry Models,” Microsoft Learn (accessed August 2026). Confirms that prompts, completions, embeddings, and training data “are NOT used to train, retrain, or improve” the base models and are not shared with other customers or with OpenAI.
https://learn.microsoft.com/en-us/azure/ai-foundry/responsible-ai/openai/data-privacy
2. Microsoft, “What is Azure Private Link?” Microsoft Learn. Private endpoints place a PaaS service on a customer VNet address and remove exposure to the public internet.
https://learn.microsoft.com/en-us/azure/private-link/private-link-overview
3. Microsoft, “Managed identities for Azure resources,” Microsoft Learn. Eliminates stored credentials for service-to-service authentication.
https://learn.microsoft.com/en-us/entra/identity/managed-identities-azure-resources/overview
4. Microsoft, “Azure Key Vault soft-delete and purge protection,” Microsoft Learn.
https://learn.microsoft.com/en-us/azure/key-vault/general/soft-delete-overview
5. GitHub, “Configuring OpenID Connect in Azure,” GitHub Docs. Federated workload identity removes long-lived cloud credentials from CI/CD.
6. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” NeurIPS 2020. arXiv:2005.11401. Establishes retrieval grounding as a method for producing more specific, diverse, and factual generation than a parametric-only model.
7. AICPA, “SOC 2® — SOC for Service Organizations: Trust Services Criteria.” The audit framework Quill’s control design is mapped to.
https://www.aicpa-cima.com/topic/audit-assurance/audit-and-assurance-greater-than-soc-2
8. U.S. Department of Health & Human Services, “Covered Entities and Business Associates.” Establishes the scope of HIPAA, which applies to human protected health information; Quill adopts equivalent controls as an engineering benchmark rather than as a veterinary legal requirement.
https://www.hhs.gov/hipaa/for-professionals/covered-entities/index.html
9. Ülkir M, Paslı B, “Reference Hallucination, Citation Reliability, and Readability of Large Language Models in Anatomy-Related Question Answering,” Clinical Anatomy, 2026. DOI: 10.1002/ca.70187.
10. Microsoft, “Hybrid search — Azure AI Search,” Microsoft Learn. Describes Reciprocal Rank Fusion over parallel vector and keyword (BM25) result sets, and semantic reranking.
https://learn.microsoft.com/en-us/azure/search/hybrid-search-overview
11. Microsoft, “Jobs in Azure Container Apps,” Microsoft Learn. Scheduled and event-driven containerized jobs with scale-to-zero.
12. Wang D, Ye J, Li J, et al., “Enhancing Large Language Models for Improved Accuracy and Safety in Medical Question Answering: Comparative Study,” Journal of Medical Internet Research, 2025. DOI: 10.2196/70190.
13. Radford et al., “Robust Speech Recognition via Large-Scale Weak Supervision” (Whisper), 2022. arXiv:2212.04356.
14. Scoresby K, Jurney C, Fackler A, Tran CV, Nugent W, Strand E, “Relationships between diversity demographics, psychological distress, and suicidal thinking in the veterinary profession: a nationwide cross-sectional study during COVID-19,” Frontiers in Veterinary Science, 2023. DOI: 10.3389/fvets.2023.1130826.
15. Rotenstein LS, Holmgren AJ, Thombley R, et al., “Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes: A Multisite Study,” JAMA, 2026. DOI: 10.1001/jama.2026.2253.
16. North American Pet Health Insurance Association (NAPHIA), “Industry Data — State of the Industry.” NAPHIA members and report participants represent 99% of policies in effect in North America; insured-pet counts and gross written premium growth are as published in the current edition.
17. ASC X12, “X12 EDI Standards” — the 837 (claim), 835 (remittance advice), 270/271 (eligibility inquiry and response) transaction sets used across insurance claim exchange.
18. HL7 International, “FHIR (Fast Healthcare Interoperability Resources) R4 Specification.”
19. PostgreSQL Global Development Group, “Row Security Policies,” PostgreSQL Documentation. Row-level security enforced by the database engine rather than by application code.
https://www.postgresql.org/docs/current/ddl-rowsecurity.html
20. Microsoft, “High availability and disaster recovery — Azure Database for PostgreSQL flexible server,” Microsoft Learn. Zone-redundant HA and point-in-time restore.
https://learn.microsoft.com/en-us/azure/postgresql/flexible-server/concepts-high-availability
21. Awawdeh L, Waran N, Forrest RH, “Aotearoa New Zealand pet guardians’ attitudes towards the financial cost of pet care,” Animal Welfare, 2025. DOI: 10.1017/awf.2025.10048.
22. Szlosek D, Coyne M, Riggott J, Knight K, McCrann DJ, Kincaid D, “Development and validation of a machine learning model for clinical wellness visit classification in cats and dogs,” Frontiers in Veterinary Science, 2024. DOI: 10.3389/fvets.2024.1348162.
Keep reading

Innovation · 13 min read
Why veterinary PIMS integrations stop short of a paid claim
Most veterinary PIMS integrations move data. A claim is only finished when it is paid. Here is how QuillsFlow is designed to close that loop.

AI/ML · 16 min read
Engineering the Intelligent Veterinary Operating System
A governed intelligence layer connecting the clinical encounter, documentation, operational decisions, claim quality, payer response, cash, and measured outcomes.