This paper addresses the challenge of evaluating large language model and embedding architecture performance in procurement-specific retrieval augmented generation systems. While existing research focuses on general-purpose benchmarks that ignore domain-specific terminology, query complexity, and real-world decision-making needs, specialized procurement applications require tailored evaluation frameworks. We evaluate three large language models (LLM) architectures, Llama3.2-3B, Phi3.5-Mini-Instruct, and Qwen3-4B-Instruct-2507, and three embedding models, EmbeddingGemma-300M, Ember-v1, and Qwen3-Embedding-0.6B, using 500 procurement-related documents and 34 domain-specific queries across 12 configurations, conducting 408 evaluations. Results show significant variation in performance, with the top-performing combination (Qwen3-4B-Instruct-2507 paired with EmbeddingGemma-300M) achieving 89.7% overall accuracy, a 30.2% improvement over the weakest configuration. We implement domain-specific fine-tuning using supervised similarity learning on 67 procurement-related terms and mean squared error (MSE) loss optimization, improving performance to 88.3%, a 2.08% relative gain across all large language model architectures. This study provides robust evidence for architectural decisions in enterprise retrieval augmented generation (RAG) systems. Our evaluation of nine LLM-embedding combinations identified the Qwen3-4B-Instruct-2507 and EmbeddingGemma-300M pair as top-performing, achieving 89.7%. Fine-tuning increased this to 91.5%, a 1.8-point absolute improvement and 2.01% relative gain, validated across all models. The findings establish new benchmarks for procurement-domain RAG systems and demonstrate that targeted fine-tuning leads to consistent performance improvements in specialized enterprise applications.