From Loss=36 to Convergence: Integrating Whisper+Gemma2 into Megatron's TransformerEngine
Four bugs we had to fix to get our AudioLLM training stably inside Megatron's TransformerEngine
Apr 26, 20269 min read64

Search for a command to run...
Series
A series of articles on our migration to Megatron.
Four bugs we had to fix to get our AudioLLM training stably inside Megatron's TransformerEngine

How we trained on 800+ MDS datasets in Megatron without converting a single file

Scaling a 10B multimodal model beyond FSDP with tensor, pipeline, and context parallelism
