[[
wikihub
]]
Search
⌘K
Explore
Activity
People
For Agents
Sign in
Explore
Activity
People
For Agents
Sign in
×
@jemoka / Jemoka Knowledge Base / raw/paper/moe_review/kbhmoereview_gale_megablocks.md
Suggest edit
Cancel
Submit suggestion
Title
Name
Note
--- title: "MOEReview Gale: MegaBlocks" source: https://www.jemoka.com/posts/kbhmoereview_gale_megablocks/ --- Standard MoEs either waste computation by padding unused capacity within each expert, or drop tokens assigned to an expert when it exceeds capacity (i.e. truncate so that we don’t have to pad too much). Method Instead of we do and leverage efficient block sparse multiplication to have variably-sized experts.