Home GPT 3 Zephyr 7B


Mixtral 8x7B


Our service is free. If you like our work and want to support us, we accept donations (Paypal).



Mistral AI

Mistral AI remains committed to providing top-notch open models for developers. Progress in AI involves exploring fresh tech avenues, not just rehashing known architectures and training methods. Crucially, it’s about sharing unique models with the community to inspire novel inventions and applications.

Mixtral 8x7B, a top-tier sparse mixture of experts model (SMoE) with open weights, licensed under Apache 2.0. It surpasses Llama 2 70B in most benchmarks, offering 6x quicker inference. It’s the most potent open-weight model with a liberal license, and the best in terms of cost-performance balance. Notably, it equals or exceeds GPT3.5 on most standard benchmarks.

Mixtral’s features include:
  - handling of a 32k token context.
  - Multilingual support for English, French, Italian, German, and Spanish.
  - High performance in code generation.
  - Can be fine-tuned into an instruction-following model, scoring 8.3 on MT-Bench.

Distributed open models

Advancing open models with sparse structures Mixtral is a sparse network of expert mixtures. It’s a decoder-only model where the feedforward block selects from 8 unique parameter groups. At each layer, for each token, a router network picks two of these groups (the ‘experts’) to process the token and add their outputs.

This method boosts the model’s parameters while managing cost and latency, as only a portion of the total parameters are used per token. Specifically, Mixtral has 46.7B total parameters but uses just 12.9B per token. Thus, it inputs and outputs at the same pace and cost as a 12.9B model.


Advertisement

Chat with Mixtral 8x7B Instruct by Mistral AI for free, without an account: a sparse mixture-of-experts model that activates two of its eight experts per token, giving the quality of a 40B-class model at the speed of a 13B one. Strong in French, English, German, Spanish and Italian, and at code.

How to use Mixtral 8x7B chat

  1. Write your question or task in the chat box, in any of the supported languages.
  2. Send it and read the streamed answer.
  3. Ask follow-up questions; the model keeps the conversation context.
  4. For code, state the language and paste the error message: Mixtral is good at debugging.

Limits and questions

What does mixture-of-experts mean?

The model contains eight expert sub-networks; a router picks two for each token. Only 13B of the 47B parameters are used at once, so it runs faster than its size suggests.

How does Mixtral compare with Llama 3 and GPT-4?

Above Llama 3 8B and GPT-3.5 on most benchmarks, below GPT-4 and Llama 3 70B. It is among the best open models for European languages.

Is my conversation private?

The chat runs through a hosted API; do not paste personal or confidential data. Conversations are not stored by this site.

Read the guide: What is Stable Diffusion?

Developed with ♥ by @arawak and @shayne for Stable Diffusion Web