Warning: I swapped our RAG pipeline for a 7B fine-tuned model and retrieval lost
Spent 6 weeks and about $800 of GPU time building a retrieval setup with a vector store over 40k docs, but a plain fine-tuned model answered our support tickets better because half our docs were stale anyway. Before you all pile on, what test would actually settle this for you?