As embedding models continue to improve, many teams are solving domain-specific NLP problems using retrieval, semantic search, and RAG pipelines instead of fine-tuning foundation models.
This raises an interesting technical question:
At what point does improving retrieval stop being enough?
For tasks involving domain expertise, specialized terminology, reasoning, or long-context understanding, when would you choose:
- Better embeddings
- Better retrieval
- Fine-tuning
- A combination of both
Have you encountered production use cases where retrieval-based approaches hit a ceiling that only fine-tuning could overcome?
Or do you think increasingly powerful embedding models are making fine-tuning less necessary for most enterprise NLP workloads?
