RAG (retrieval-augmented generation) is an architecture that grounds large language model responses in a private knowledge base by retrieving relevant passages at query time and passing them to the LLM as context. It eliminates the "knowledge cutoff" problem, enables per-tenant data isolation and lets organizations use LLMs on proprietary documents without fine-tuning. Production RAG requires a document ingestion pipeline, embedding model, vector store, retrieval strategy, reranking, citation tracking, evaluation harness and drift monitoring. Pharos has shipped RAG systems for legal research, customer support, internal documentation, healthcare Q&A and technical support since 2023.
Authoritative citations
12 sources
-
Stanford AI Index
The Stanford AI Index tracks multi-year movement on ML benchmarks, training compute, responsible AI metrics and enterprise adoption across industries, making it the most cited yearly reference for grounding ML investment cases.
aiindex.stanford.edu
-
Papers With Code
Papers With Code maintains live state-of-the-art leaderboards for ML tasks across image classification, object detection, NLP and tabular prediction, which we use to pick baselines before committing to a model family.
paperswithcode.com
-
arXiv, Chen and Guestrin 2016
The XGBoost paper by Chen and Guestrin remains the most cited gradient boosting reference and underpins tabular ML baselines we still ship in FinTech and logistics systems a decade after publication.
arxiv.org
-
arXiv, LightGBM
Microsoft Research LightGBM introduced leaf-wise tree growth and histogram-based splits, giving lower latency and memory footprint than XGBoost on wide tabular data, which is why our fraud detection stack defaults to it.
arxiv.org
-
McKinsey State of AI
McKinsey documents annual enterprise ML adoption across functions like marketing, service operations and supply chain, and consistently reports that scaled ML correlates with higher EBIT contribution versus pilot-only organizations.
mckinsey.com
-
Gartner AI Hype Cycle
Gartner maps enterprise ML techniques across the hype cycle phases, flagging which capabilities are production-ready for mid-market adoption versus still speculative, which we cross-check before recommending a build path.
gartner.com
-
IDC Worldwide AI Spending Guide
IDC publishes the worldwide AI spending guide with multi-year forecasts by industry, use case and geography, which we reference when sizing three-year total cost of ownership for ML platform engagements.
idc.com
-
NIST AI Risk Management Framework
The NIST AI RMF defines a govern, map, measure and manage lifecycle for AI systems that we apply to production ML including model cards, bias testing and incident response procedures for regulated deployments.
nist.gov
-
OWASP ML Security Top 10
OWASP maintains a ranked list of the top machine learning security risks including input manipulation, training data poisoning, model theft and adversarial attacks, which we use as a threat model checklist before exposing any ML endpoint.
owasp.org
-
O'Reilly AI Adoption in the Enterprise
The O'Reilly AI adoption survey tracks ML maturity stages across enterprises, reporting on deployment percentages, skills gaps and the most common production blockers which consistently include data quality and monitoring rather than model choice.
oreilly.com
2022
-
Google Cloud MLOps Architecture
Google Research published the canonical MLOps continuous delivery reference describing three maturity levels from manual to fully automated pipelines, which we use as the template for client MLOps roadmaps and capability gap assessments.
cloud.google.com
-
PyTorch Blog
The PyTorch engineering blog tracks the 2.x production tooling surface including torch.compile, TorchServe updates and quantization workflows, which shape our default serving stack for sub-50ms p99 inference on GPU and CPU targets.
pytorch.org
What we do not do
- Use cases where a traditional search index (Elasticsearch, Typesense) delivers better relevance at lower cost
- Knowledge bases with fewer than 100 documents (few-shot prompting is simpler)
- Real-time chat where 3-second retrieval latency is too slow
- Projects without a labeled evaluation set for retrieval quality