AI INFO FORGE
Model Releases The Verge AI Jul 27, 2026

Why China is giving away its best AI models

Chinese startup Moonshot released Kimi K3, an open-weight AI model that reportedly matches or exceeds performance of leading US models while being significantly cheaper to develop. The free release of model weights to developers has raised concerns in the US tech industry about maintaining competitive advantages through proprietary systems.

Read on The Verge AI →
Model Releases TechCrunch AI Jul 24, 2026

Anthropic launches Opus 5

Anthropic has released Opus 5, a new AI model that offers improved cost efficiency and fewer operational restrictions compared to its predecessor Fable, potentially making it more attractive for widespread applications.

Read on TechCrunch AI →
Model Releases The Verge AI Jul 24, 2026

Meta is making its AI chatbot more like an assistant

Meta has upgraded its AI chatbot with new productivity capabilities, including calendar integration for event planning, daily briefings, and interactive research features. The company deployed its Muse Spark 1.1 model to enhance the assistant's functionality as it competes with ChatGPT, Gemini, and Claude.

Read on The Verge AI →
Model Releases The Verge AI Jul 23, 2026

Claude’s voice mode is now available for Opus and Sonnet

Anthropic has expanded voice mode access to its more capable Claude Opus and Sonnet models, moving beyond the previously exclusive Haiku model. The expansion includes integration with productivity apps like Gmail, Slack, and Canva, enabling users to tackle complex business problems through voice interaction.

Read on The Verge AI →
Model Releases arXiv cs.AI Jul 22, 2026

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio

Researchers introduced Athena-Brain-8B, an 8-billion parameter language model optimized as an on-device controller for robots. The model combines general language understanding with specialized embodied interaction capabilities through multi-stage training, achieving competitive performance on both general benchmarks and robot-specific tasks while producing concise responses for efficient hardware execution.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 21, 2026

FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images

Researchers introduced FUSAR-R1, a reasoning-focused vision-language model designed specifically for Synthetic Aperture Radar image interpretation. The model uses chain-of-thought reasoning data and reinforcement learning to perform expert-level analysis tasks like target detection, counting, and land-cover classification, outperforming existing multimodal models on SAR-specific challenges.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 21, 2026

Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation

Researchers introduced Pailitao-MMSearch, a specialized multimodal search model for e-commerce that handles text, images, and voice queries simultaneously. Built on Qwen and deployed on Taobao, the model achieved significant improvements in online metrics, including 13.61% increases in merchandise volume compared to traditional multimodal search approaches.

Read on arXiv cs.AI →
Model Releases Hugging Face Jul 20, 2026

Introducing Cosmos 3 Edge

Hugging Face announced Cosmos 3 Edge, a new model or tool in their Cosmos product line. The specific capabilities and technical details were not provided in the available information.

Read on Hugging Face →
Model Releases The Verge AI Jul 20, 2026

China delivers a one-two punch to America’s AI dominance

Chinese AI companies Moonshot and Alibaba released new models claiming competitive performance with leading U.S. systems like OpenAI and Anthropic, at lower costs. The releases signal narrowing technological gaps between Chinese and American AI capabilities amid growing geopolitical competition.

Read on The Verge AI →
Model Releases arXiv cs.AI Jul 20, 2026

Cura 1T: Specialized Model for Agentic Healthcare

Researchers introduced Cura 1T, a specialized language model designed for healthcare tasks including patient consultation, clinical reasoning, and electronic health record tool use. The model was developed using a human-gated self-evolution training approach that iteratively improves capabilities by targeting observed failures with synthetic and curated data, achieving competitive performance on healthcare and general reasoning benchmarks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 20, 2026

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

Researchers introduced S1-Omni, a multimodal AI model designed to handle diverse scientific tasks including molecular generation, protein structure prediction, and spectrum analysis. The unified model integrates scientific laws and expert knowledge while processing multiple data types, demonstrating competitive performance against specialized domain models and general-purpose AI systems on 60+ scientific benchmarks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 20, 2026

Loop the Loopies!

Researchers introduced Loopie, a series of looped Transformer models using Mixture-of-Experts architecture that achieves competitive performance compared to larger vanilla Transformers trained with equivalent compute resources. The models demonstrate strong reasoning capabilities, achieving gold-medal results on the 2025 IMO and IPhO competitions.

Read on arXiv cs.AI →
Model Releases TechCrunch AI Jul 18, 2026

Kimi: Threat or menace?

Moonshot AI, a Chinese company, released an updated version of its Kimi model, sparking debate about potential implications for AI development and geopolitical considerations.

Read on TechCrunch AI →
Model Releases Google DeepMind Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google unveiled Gemini 3.5 Flash Cyber, a specialized lightweight AI model designed to identify and remediate security vulnerabilities. The model combines efficiency with cybersecurity capabilities for vulnerability detection and patching.

Read on Google DeepMind →
Model Releases AWS Machine Learning Jul 16, 2026

Introducing Grok on Amazon Bedrock

xAI's Grok 4.3 model is now available on Amazon Bedrock, enabling enterprises to build AI agents and workflows with a model featuring configurable reasoning, tool-use capabilities, and a 1 million token context window. The integration positions xAI as a new model provider on Bedrock's inference platform.

Read on AWS Machine Learning →
Model Releases arXiv cs.AI Jul 16, 2026

EZSMT Version 3, Matured

Researchers introduced EZSMTV3, an enhanced framework for Constraint Answer Set Programming that combines logic programming with constraint solving. The system leverages existing SMT solvers and demonstrates improved language expressiveness, optimization capabilities, and support for mixed-domain constraints compared to competing CASP systems.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 16, 2026

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Researchers released Boogu-Image-0.1, an open-source multimodal model family that performs text-to-image generation, editing, and bilingual text rendering. The model achieves performance comparable to closed-source alternatives while trained on 208.62 million images with approximately $400K computational cost, with code and weights made publicly available.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 16, 2026

OvisOCR2 Technical Report

OvisOCR2 is a 0.8B document parsing model that converts document images to Markdown format covering text, formulas, tables, and visual elements. Trained using a combination of real and synthetic data with supervised fine-tuning and reinforcement learning, it achieves state-of-the-art performance on multiple benchmarks, surpassing pipeline-based methods in end-to-end document parsing.

Read on arXiv cs.AI →
Model Releases Wired AI Jul 15, 2026

Thinking Machines Lab Drops Its First Model

Thinking Machines Lab released Inkling, an open-source model with 975 billion parameters designed to process both video and audio. The release positions the startup as a competitor in the generative AI space alongside established players like Anthropic and OpenAI.

Read on Wired AI →
Model Releases arXiv cs.AI Jul 14, 2026

QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics

Alibaba researchers introduced QwenPaw-Data, an autonomous agent system designed for enterprise data analysis that integrates data warehouses, dashboards, and documents to convert natural language requests into analytical workflows. The system uses interconnected metadata graphs, reusable analytical skills, and controlled execution to improve data access reliability and analytical quality while continuously learning from feedback.

Read on arXiv cs.AI →
Model Releases The Verge AI Jul 13, 2026

Siri AI is already changing how I use my iPhone

Apple released the first public beta of iOS 27, which prioritizes performance improvements and bug fixes over new features. Key updates include faster app launches and search, enhanced Messages functionality with RCS encryption, and refinements to Liquid Glass technology.

Read on The Verge AI →
Model Releases arXiv cs.AI Jul 13, 2026

A Sovereign, Open-Source Foundation Model for German and English

Researchers released Soofi S 30B-A3B, an open-source hybrid language model optimized for German and English that uses a Mixture-of-Experts architecture to activate only 3B of its 30B parameters per token. Trained on 27 trillion tokens with emphasis on German, it matches larger dense models on benchmarks while achieving superior code performance and outperforming other European sovereign models, with all weights and training artifacts to be publicly released.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 13, 2026

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

Researchers introduced ALICE, a unified foundation model for pathology that combines knowledge from eight specialized models through multi-stage distillation. Trained on nearly 25 million pathology images, ALICE outperformed task-specific models across 21 evaluation scenarios and 96 downstream tasks in tissue analysis, vision-language tasks, and whole-slide assessment.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 10, 2026

Infinity-Parser2 Technical Report

Researchers introduced Infinity-Parser2, a multimodal model for document parsing that uses synthetic data generation and multi-task reinforcement learning. The team released a 5-million-sample bilingual dataset and two model variants—Flash for speed and Pro for accuracy—achieving state-of-the-art results on multiple parsing benchmarks.

Read on arXiv cs.AI →
Model Releases The Verge AI Jul 8, 2026

ChatGPT’s upgraded voice mode is better at shutting up

OpenAI unveiled GPT-Live-1, an upgraded voice model for ChatGPT that interrupts users less frequently and better simulates natural conversation. The model can delegate tasks to more capable text models like GPT-5.5 for complex reasoning or web searches, improving response quality and speed.

Read on The Verge AI →
Model Releases arXiv cs.AI Jul 8, 2026

KAT-Coder-V2.5 Technical Report

KAT-Coder-V2.5 is an autonomous coding agent trained to work directly within real code repositories rather than generating isolated snippets. The model employs advanced post-training techniques including sandboxed environment reconstruction, reinforcement learning optimization, and multi-teacher distillation, achieving competitive performance on software engineering benchmarks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 8, 2026

Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context

Harrison.Rad 1.5 is a radiology-focused AI model that generates medical reports by analyzing X-ray images alongside patient history and clinical context. Trained through domain adaptation and vision-language techniques, it achieved the highest performance on clinical benchmarks and meets standards equivalent to professional radiology certification exams.

Read on arXiv cs.AI →
Model Releases OpenAI Jul 8, 2026

Introducing GPT-Live

OpenAI has launched GPT-Live, an updated voice model generation designed to enable more natural conversations between humans and AI systems. The technology now powers ChatGPT's voice feature.

Read on OpenAI →
Model Releases The Verge AI Jul 7, 2026

Anthropic is launching Claude Cowork on mobile and web

Anthropic is expanding Claude Cowork, its collaborative AI platform, to mobile and web interfaces starting Tuesday. Previously limited to desktop apps, the feature will roll out first to Max subscribers, with full functionality remaining on desktop while mobile and web versions offer core collaboration capabilities.

Read on The Verge AI →
Model Releases arXiv cs.AI Jul 7, 2026

iFLYTEK-Embodied-Omni Technical Report

iFLYTEK released Embodied-Omni, a unified multimodal foundation model that integrates vision, language, and action processing in a single framework for robotic control tasks. The model uses a brain-cerebellum architecture where vision-language components handle high-level planning while action components directly execute instructions, trained on both human demonstrations and robot interaction data.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 7, 2026

Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine

Researchers introduced OpenDDE, an open-source AI model for drug discovery that predicts biomolecular structures and relationships using co-folding technology. The system integrates structure prediction with additional capabilities for drug design and optimization, with released code and benchmarks intended to advance collaborative research in computational biology.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 7, 2026

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

Nvidia researchers introduced Nemotron-Labs-3-Puzzle-75B-A9B, a compressed version of their Nemotron-3-Super model designed for efficient deployment. Using a multi-stage compression pipeline combining pruning, distillation, and quantization, the model achieves 2x higher server throughput on interactive workloads and enables 8x more concurrent long-context requests on a single H100 GPU while maintaining competitive performance across reasoning, coding, and multilingual benchmarks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 7, 2026

Gemma 4 Technical Report

Google introduced Gemma 4, an open-weight multimodal language model family ranging from 2.3B to 31B parameters. The models feature improved vision and audio processing, a reasoning mode for generating explanations, and optimizations for inference efficiency and long-context understanding, achieving performance competitive with larger models on various benchmarks.

Read on arXiv cs.AI →
Model Releases AWS Machine Learning Jul 6, 2026

Teaching models to forget: Selective unlearning with Amazon Nova

AWS introduced Reverse Direct Preference Optimization (rDPO), a technique that allows organizations to selectively reduce content moderation restrictions in Amazon Nova models for legitimate business purposes. The Customizable Content Moderation Settings feature lets approved customers adjust safeguards across safety, and other responsible AI categories while maintaining model quality.

Read on AWS Machine Learning →
Model Releases arXiv cs.AI Jul 3, 2026

The Wiola Architecture for Efficient Small Language Models

Researchers introduced Wiola, a novel small language model architecture featuring five new components: spiral rotary positional encoding, gated cross-layer attention, adaptive token merging, dual stream feed-forward networks, and modified normalization. The architecture was released in four sizes ranging from 120M to 1.5B parameters and is compatible with HuggingFace Transformers.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 2, 2026

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Seed2.0 is a new model series designed to handle complex real-world tasks by addressing long-tail knowledge gaps and improving instruction-following capabilities. The developers built a specialized evaluation framework based on genuine user needs and realistic scenarios to guide the model's development, achieving improvements in reasoning, visual understanding, and search functionality.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jul 1, 2026

Xiaomi-GUI-0 Technical Report

Xiaomi introduced Xiaomi-GUI-0, a GUI agent trained and evaluated on real mobile devices rather than simulated environments, addressing the gap between benchmark performance and real-world usability. The system uses a hybrid infrastructure combining physical devices with sandboxes and employs multi-source training data with an error-driven feedback loop to improve stability across real applications.

Read on arXiv cs.AI →
Model Releases AWS Machine Learning Jul 1, 2026

Safely Releasing Frontier Models to Customers

AWS announced that Anthropic's Claude Fable 5 models will be available on Amazon Bedrock with enhanced safety features. AWS emphasized its commitment to secure model deployment while balancing rapid access to frontier AI capabilities for customers against broader societal security considerations.

Read on AWS Machine Learning →
Model Releases AWS Machine Learning Jun 30, 2026

Introducing Claude Sonnet 5 on AWS: Anthropic’s most capable Sonnet model

Anthropic released Claude Sonnet 5 on Amazon Bedrock and Claude Platform on AWS, positioning it as the most capable Sonnet model with improved performance for coding and agentic tasks while maintaining cost efficiency. The model integrates with AWS infrastructure, offering enterprise security, regional data residency, and unified billing alongside Anthropic's native platform features.

Read on AWS Machine Learning →
Model Releases The Verge AI Jun 28, 2026

China’s Z.ai claims it can match Mythos on cybersecurity

Zhipu AI released GLM-5.2, an open-weight model that researchers claim matches performance with advanced security-focused models in bug-finding and cybersecurity tasks. Though the model trails US competitors on general benchmarks, the release signals China's narrowing capability gap, raising concerns for the US government's ongoing efforts to restrict China's access to advanced AI systems.

Read on The Verge AI →
Model Releases arXiv cs.AI Jun 27, 2026

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

Researchers introduced ReasonCLIP-58M, an enhanced version of CLIP that incorporates reasoning supervision through a two-stage training approach. The framework includes new datasets and a benchmark to improve visual reasoning capabilities, and shows performance gains when integrated into multimodal systems like LLaVA-NeXT without increasing inference costs.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 27, 2026

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Researchers introduced Wan-Streamer, a foundation model built for real-time audio-visual interaction with approximately 200ms response latency. Unlike traditional systems using separate modules for speech recognition, language processing, and video generation, this model integrates all capabilities in a single Transformer architecture with specialized streaming mechanisms.

Read on arXiv cs.AI →
Model Releases The Verge AI Jun 26, 2026

OpenAI unveils GPT-5.6 amid US AI regulatory drama

OpenAI released GPT-5.6, a new model suite with three variants (Sol, Terra, Luna) following a staggered release arrangement requested by the Trump administration. The flagship Sol model emphasizes capabilities in coding, cybersecurity, and biology, with pricing at $5/$30 per million tokens.

Read on The Verge AI →
Model Releases arXiv cs.AI Jun 26, 2026

Discovering Millions of Interpretable Features with Sparse Autoencoders

Researchers released Qwen3-Instruct SAE, a collection of sparse autoencoders trained on Qwen3 language models to extract interpretable features from neural representations. The work demonstrates how these SAEs can identify and manipulate specific behaviors, such as refusal responses, providing tools for understanding and steering model behavior.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 25, 2026

ZONOS2 Technical Report

ZONOS2 8B, a new text-to-speech model, scales up from its predecessor to 8 billion parameters using a mixture-of-experts architecture and trains on over 6 million hours of audio data. The model achieves competitive performance on naturalness, prosody, and voice cloning while maintaining efficient streaming latency, with weights and code released publicly.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 25, 2026

FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation

Researchers introduced FISHER, a foundation model designed for analyzing industrial signals across multiple data types and sampling rates. The model uses a novel sub-band approach to handle varying sampling rates without resampling and is pre-trained using self-distillation on audio data. FISHER outperforms larger specialized models while being significantly smaller, with researchers releasing both the model and a new 19-dataset benchmark.

Read on arXiv cs.AI →
Model Releases The Verge AI Jun 24, 2026

OpenAI reveals its first AI processor: Jalapeño

OpenAI unveiled Jalapeño, a custom AI processor chip developed with Broadcom designed specifically for AI inference tasks in servers. The ASIC chip is intended to support both current and future large language models, marking OpenAI's entry into custom hardware for AI operations.

Read on The Verge AI →
Model Releases Wired AI Jun 20, 2026

Siri AI Hands On: A Smart, Helpful Assistant

Apple's updated Siri AI features improved conversational abilities and expanded availability across devices, with enhanced practical utility for users. The assistant demonstrates stronger natural language understanding and responsiveness compared to previous iterations.

Read on Wired AI →
Model Releases arXiv cs.AI Jun 20, 2026

IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources

Researchers introduced IHUBERT, a Persian language model trained on 7-8 billion tokens using semantic deduplication and domain-balanced preprocessing to address data scarcity. The 125M-parameter model achieved top performance on Persian question answering and natural language inference tasks while remaining competitive on other NLU benchmarks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 20, 2026

SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

Researchers introduced SleepMaMi, a foundation model designed for sleep medicine that analyzes both full-night sleep patterns and detailed biosignal features from polysomnography recordings. Trained on over 20,000 PSG recordings, the model uses dual encoders and demographic-guided learning to outperform existing approaches across multiple clinical sleep analysis tasks.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 20, 2026

TerraMind: Large-Scale Generative Multimodality for Earth Observation

TerraMind is a new multimodal foundation model designed for Earth observation that processes data at both token and pixel levels across nine geospatial data types. The model demonstrates strong performance on standard benchmarks and introduces a capability to generate synthetic data during training and inference, with researchers releasing the model weights and code openly.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 20, 2026

Vero: An Open RL Recipe for General Visual Reasoning

Researchers introduced Vero, an open-source family of vision-language models trained with reinforcement learning across diverse visual reasoning tasks. The framework combines 600K samples from 59 datasets with task-specific reward mechanisms, enabling models to match or exceed existing open-weight competitors on benchmarks spanning charts, science, and spatial understanding.

Read on arXiv cs.AI →
Model Releases arXiv cs.AI Jun 19, 2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek released V4 series language models featuring two MoE variants (Pro with 1.6T parameters and Flash with 284B parameters) that support one million token contexts. The models incorporate architectural improvements including hybrid attention mechanisms and achieve significantly improved efficiency, requiring 27% fewer inference FLOPs and 10% of KV cache compared to the previous V3.2 version.

Read on arXiv cs.AI →
Scott Sokolowski, founder of AI Info Forge

Hi — I'm Scott. I'm building AI Info Forge because the Claude skills market is full of prompts that worked last month and break this month. Verified, drift-tested skills is the version of this market I wish existed. If you've felt the same friction, get on the list and I'll send you early access.

— Scott Sokolowski, Founder