Mixtral contains several expert blocks, while a router activates only two per layer and token. This increases total capacity without proportional inference cost; open weights and the Apache 2.0 licence allowed researchers and companies to deploy and adapt the system themselves.
Source 1: Mistral AI · техническое описаниеSource 2: Mistral AI · карточка моделиMixtral 8×7B
About this eventAn open-weight sparse mixture of experts in which only part of the parameters runs for each token.