Enterprise RAG Knowledge Assistant
An internal AI assistant that lets clinical and administrative staff at a regional hospital network search policies, SOPs, and internal documentation in plain language.
A regional hospital network with thousands of policy documents, SOPs, and internal clinical guidelines spread across shared drives and a legacy intranet.

The right policy existed — nobody could find it fast enough
Clinical and administrative staff needed answers to questions like "what's the current infection-control protocol for this unit?" or "what's the escalation path for a medication discrepancy?" — and the correct answer usually existed somewhere in a PDF, a wiki page, or an old intranet post.
Search across those systems was keyword-only and inconsistent. Staff either called someone who might know, or used whatever version of the document they could find first — which was sometimes outdated. On-call and night-shift staff had the least support, since the people who usually answered these questions were asleep.
- Answers existed but were scattered across shared drives, an intranet, and PDFs with no unified search
- No way to know if a found document was the current version
- Time-sensitive questions (infection control, escalation paths) needed faster answers than a phone tree could give
- Any answer had to be traceable back to its source document — staff won't act on an unsourced answer
Ask in plain language, get an answer with its source attached
TecXra built a retrieval-augmented generation (RAG) assistant that indexes the hospital network's policy library, SOPs, and internal knowledge base into a vector database. Staff ask a question in plain language; the system retrieves the most relevant passages, reranks them for accuracy, and generates an answer grounded in — and citing — the actual source document.
Nothing is answered from the model's general knowledge alone. If the retrieval step doesn't find a confident match, the assistant says so instead of guessing, and points the user toward who to ask.
A look inside the workflow
Every document goes through the same pipeline before it's searchable, and every answer traces back through that same pipeline to its source.
Documents
Policies, SOPs, and guidelines ingested from shared drives and the intranet.
Chunking
Documents are split into retrieval-sized passages that preserve context.
Embeddings
Each passage is converted into a vector representation of its meaning.
Vector Database
Embeddings are indexed for fast semantic — not just keyword — search.
Retrieval
A staff question is embedded and matched against the closest passages.
Reranking
Candidate passages are reordered for relevance before reaching the model.
LLM
Generates an answer grounded strictly in the retrieved passages.
Answer
Returned to staff with the source document cited inline.
What the system actually does
Semantic search, not keyword search
Finds the right passage even when staff don't use the document's exact wording.
Every answer is sourced
Responses cite the specific document and section they came from.
Versioning-aware indexing
Superseded documents are flagged so staff aren't served outdated policy.
Scoped access
Role-based access ensures staff only retrieve documents they're permitted to see.
On-call friendly
Available around the clock for staff who don't have a supervisor to ask at 3am.
Says "I don't know" on purpose
Low-confidence retrieval returns a clear "not found" instead of an unsupported guess.
Built on
AI / Retrieval
Orchestration
Data
Infrastructure
What it looks like in use


Every passage is chunked, embedded, and indexed the same way — so every answer traces back through the same pipeline to its source.
What changed
Illustrative Impact — representative outcomes for this class of system, not measured client figures.
Faster answers
Staff get a sourced answer in seconds instead of searching multiple systems or waiting on a callback.
Fewer outdated-document mistakes
Version-aware indexing reduces the chance of acting on a superseded policy.
On-call support improved
Night-shift and on-call staff get the same answer quality as daytime staff.
Traceable, auditable answers
Every response can be checked against its cited source document.
How it was built
Document audit & ingestion
Catalogued and ingested the policy library, resolving duplicate and outdated versions.
Chunking & embedding pipeline
Built the pipeline that keeps the vector index current as documents change.
Retrieval tuning
Tuned chunking and reranking against real staff questions for retrieval accuracy.
Access control & rollout
Layered in role-based access and rolled out by department, starting with the highest-volume queries.
Buried in internal documentation?
If your team already has the right answers written down somewhere, let's talk about making them instantly findable.
