The developer, who began the project six months ago while working with 15,000 hours of audio data stored in Google Cloud Storage, faced a critical constraint: the dataset was too large to transfer and process locally. Rather than accept the high bandwidth and compute costs of cloud-based training, they built a specialized fine-tuner for Apple's M2 Ultra Mac Studio with explicit attention to memory efficiency and limited compute budgets. The resulting tool enables developers to perform multimodal model training—processing both audio and text simultaneously—on Apple Silicon without relying on expensive GPU clusters or cloud services like AWS or Google Cloud. This represents a meaningful shift in accessibility for machine learning practitioners who previously had to choose between cloud costs or abandoning their projects entirely.