Highlights
- Pro
Popular repositories Loading
-
Megatron-LM
Megatron-LM PublicForked from swiss-ai/Megatron-LM
Two-expert modality-routed MoE output head for Apertus: a learned top-1 router routes text tokens to a text expert and vision/audio to a VA expert (aux-loss trained). EP=1/EP=2, 1:1 modality interl…
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.
