Introducing Gemini 3.7 Flash
Gemini 3.7 Flash is our most intelligent workhorse model yet for coding and agents.
More from the blog
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
A Blog post by Ai2 on Hugging Face
Putting sign language AI into users’ hands
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Gemini Robotics 2 brings whole body intelligence to robots
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
Related papers
A comparative review of modern large language model paradigms: GPT-4, BERT, Gemini, and DeepSeek
Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).
LLMs Leak Training Data Beyond Verbatim Memorization: Extraction via Membership Decoding
The Membership Decoding method is a plug-and-play replacement for standard decoding that requires only black-box token probabilities, and a new token-level membership inference method is proposed by leveraging likelihood from reference models, shifting the generation from the original token distribution to the member token distribution.
A Multiagent Large Language Model–Based System for Early-Stage Building Layout Planning
A multiagent large language model (LLM)–based system for early-stage building layout planning, which enables flexible design requirement inputs and robust spatial reasoning and demonstrated significant improvements in both geometric quality and semantic alignment over a baseline LLM-only system.