A model architecture that routes each token to a subset of specialist sub-networks. Increases total parameters without proportionally increasing inference cost.
architecture
A model architecture that routes each token to a subset of specialist sub-networks. Increases total parameters without proportionally increasing inference cost.