AI Training Model
Gemma
What is Gemma?
Gemma is a family of lightweight, advanced, open AI models developed by Google DeepMind and other teams at Google. Based on the same technology as the Gemini models, it aims to help developers and researchers build responsible AI applications. The Gemma model family includes models with two weight scales: Gemma 2B and Gemma 7B, offering pre-trained and instruction-based fine-tuning versions, and supporting multiple frameworks such as JAX, PyTorch, and TensorFlow for efficient operation on various devices. The second-generation model, Gemma 2, was released on June 28th.
Gemma’s official website
- Gemma’s official website homepage: https://ai.google.dev/gemma?hl=zh-cn
- Gemma’s Hugging Face model: https://huggingface.co/models?search=google/gemma
- Gemma’s Kaggle model address: https://www.kaggle.com/models/google/gemma/code/
- Gemma’s technical report: https://storage.googleapis.com/deepmind-media/gemma/gemma-report.pdf
- The official PyTorch implementation is available on GitHub: https://github.com/google/gemma_pytorch
- The Google Colab installation directory for Gemma is: https://colab.research.google.com/github/google/generative-ai-docs/blob/main/site/en/gemma/docs/lora_tuning.ipynb
Key features of Gemma
- Lightweight architecture : The Gemma model is designed to be lightweight, making it easy to run in a variety of computing environments, including PCs and workstations.
- Open Model : The weights of the Gemma model are open, allowing users to use and distribute it commercially while adhering to licensing agreements.
- Pre-training and instructional fine-tuning : Provides a pre-trained model and an instructionally fine-tuned version, the latter using human feedback reinforcement learning (RLHF) to ensure the responsible behavior of the model.
- Multi-framework support : Gemma supports major AI frameworks such as JAX, PyTorch, and TensorFlow, and provides a toolchain through Keras 3.0 to simplify the inference and supervised fine-tuning (SFT) process.
- Security and Reliability : In its design, Gemma followed Google’s AI principles, using automated technology to filter sensitive information in the training data and conducting a series of security assessments, including red team testing and adversarial testing.
- Performance optimization : The Gemma model is optimized for hardware platforms such as NVIDIA GPUs and Google Cloud TPUs to ensure high performance across different devices.
- Community support : Google provides free resources on platforms such as Kaggle and Colab, as well as credits on Google Cloud, to encourage developers and researchers to use Gemma for innovation and research.
- Cross-platform compatibility : Gemma models can run on a variety of devices, including laptops, desktops, IoT devices, and the cloud, supporting a wide range of AI capabilities.
- Responsible AI Toolkit : Google also released the Responsible Generative AI Toolkit to help developers build safe and responsible AI applications, including a safety classifier, debugging tools, and application guidelines.
Gemma’s key technical points
- Model Architecture : Gemma is built on a Transformer decoder, one of the most advanced model architectures in Natural Language Processing (NLP). It employs a multi-head attention mechanism, allowing the model to focus on multiple parts of the text simultaneously. Furthermore, Gemma uses Rotated Position Embedding (RoPE) instead of Absolute Position Embedding to reduce model size and improve efficiency. The GeGLU activation function replaces the standard ReLU non-linear activation, and normalization is applied to the input and output of each Transformer sub-layer.
- Training Infrastructure : Gemma models are trained on Google’s TPUv5e, a high-performance computing platform designed specifically for machine learning. By sharding models and replicating data across multiple Pods (chip clusters), Gemma efficiently utilizes distributed computing resources.
- Pre-training data : The Gemma models are pre-trained on a large amount of English data (the 2B model is pre-trained on approximately 2 trillion tokens, while the 7B model is based on 6 trillion tokens), primarily from online documents, mathematical data, and code. The pre-training data is filtered to reduce unwanted or unsafe content while ensuring data diversity and quality.
- Fine-tuning strategy : The Gemma model is fine-tuned through supervised fine-tuning (SFT) and reinforcement learning based on human feedback (RLHF). This includes using synthetic text pairs and human-generated cue-response pairs, as well as a reward model trained on human preference data.
- Safety and Responsibility : Gemma was designed with model safety and responsibility in mind, including filtering data during the pre-training phase to reduce the risk of sensitive information and harmful content. Furthermore, Gemma has undergone a series of safety assessments, including automated benchmarking and human evaluation, to ensure the model’s safety in real-world applications.
- Performance Evaluation : Gemma underwent extensive performance evaluations across multiple domains, including question answering, commonsense reasoning, mathematical and scientific problem solving, and coding tasks. The Gemma model was compared to open models of similar or larger scale, outperforming models such as Llama-13B or Mistral-7B in 11 out of 18 benchmark tests, including MMLU and MBPP.
- Openness and accessibility : Gemma models are released as open source, providing pre-trained and fine-tuned checkpoints, as well as open-source code libraries for inference and deployment. This enables researchers and developers to access and leverage these advanced language models, driving innovation in the field of AI.