Identify the missing capability
If an assistant cannot answer questions about an updated internal handbook, the immediate problem may be access to information. If it receives the right information but repeatedly produces the wrong structure, the issue may be task behavior. These failures deserve different experiments.
Understand retrieval
Retrieval-augmented generation combines retrieved information with generation. The original RAG paper describes a model that combines parametric and non-parametric memory. In an application, retrieval quality and the quality of the final answer should be assessed separately.
Compare practical options
Begin with a clear instruction and representative examples. For changing documents, test retrieval with source references and access controls. Consider fine-tuning only with suitable training examples and a measurable behavior target. Retrieval and fine-tuning can be combined, but combining them also introduces more components to maintain.
Design a fair comparison
Use the same held-out questions for each approach. Track answer quality, unsupported claims, update effort, latency, and cost. Include questions that cannot be answered from the approved material. A convincing demonstration should show where the system refuses or asks for clarification, not only its best answers.
Try this: Classify ten observed failures as missing information, poor retrieval, instruction failure, or unsupported reasoning before deciding to train anything.