Our San Francisco fine-tuning methodology is engineered for the technical sophistication and velocity that the market demands. We operate on what we call the Production Fine-Tuning Framework -- a systematic approach that treats fine-tuning as a production engineering discipline rather than a research experiment. Phase 1: Baseline and Evaluation Architecture (Week 1) -- before any training, we establish rigorous evaluation. We build a held-out test set with the client's domain experts, define quantitative metrics aligned with business outcomes (not just perplexity or BLEU scores), and benchmark the current baseline (whether that is GPT-4 API, an existing fine-tuned model, or human performance). This evaluation architecture persists throughout the project, ensuring that every training decision is evidence-based. Phase 2: Data Engineering (Weeks 1-2) -- training data quality determines fine-tuning outcomes more than any other factor. We curate, clean, and format training data with attention to the specific patterns that drive task performance. For cost-optimization projects, this means identifying the minimum effective training set -- the smallest dataset that achieves the target accuracy on a smaller model. For capability-specialization projects, this means assembling diverse examples that cover the full range of edge cases the model will encounter in production. Phase 3: Training and Iteration (Weeks 2-3) -- we train using current best practices: LoRA/QLoRA for parameter-efficient fine-tuning where appropriate, full fine-tuning for deep domain adaptation, and multi-task training for models that must handle several related tasks. Training is iterative -- we evaluate after each run, analyze failure modes, adjust training data or hyperparameters, and retrain. SF companies expect this iteration to happen at engineering velocity, not research timelines. Phase 4: Production Deployment (Weeks 3-4) -- fine-tuned models are deployed with production-grade infrastructure: model serving optimized for the client's latency and throughput requirements (vLLM, TensorRT-LLM, or other serving frameworks), monitoring for accuracy degradation, data drift, and usage patterns, and A/B testing infrastructure to validate fine-tuned model performance against the production baseline. We do not consider a fine-tuning project complete until the model is serving production traffic and demonstrating measured improvement over the baseline. For cost optimization specifically -- the most common SF use case -- our approach is systematic. We benchmark the current API-based system's accuracy and cost, identify the smallest model that can achieve target accuracy after fine-tuning (typically 7B-13B for many tasks), fine-tune with aggressive evaluation, and deploy on cost-optimized inference infrastructure. The typical outcome is 70-90% inference cost reduction with accuracy parity or improvement.