← ALL NEWS

AI NEWS (SMOL.AI) · 13 Aug 2026

Gemini 3.7 Flash, DeepSeek Harness, and Ultrafast Inference Updates

Recent artificial intelligence updates feature major model releases, specialized inference hardware, and open-source agent tools. Google launched Gemini 3.7 Flash, a mid-tier model with significant coding and agent benchmark gains, accompanied by an introductory 50 percent price cut through the end of the year. OpenAI previewed Ultrafast mode for GPT-5.6 Sol, powered by Cerebras, achieving speeds up to 750 tokens per second to address latency in real-time workflows.

Open-source and model releases expanded across the industry. DeepSeek released DeepSeek-V4-Pro on Hugging Face alongside DeepSeek Harness, an open-source agent runtime designed with a plugin architecture for long-running workflows. MiniMax introduced MiniMax-Music3, an open-weights music generation model, and updated its video editing system. In the open-weights community, attention focused on large-scale releases like Qwen 3.8, alongside discussions regarding the immense hardware requirements needed to run multi-trillion-parameter Mixture-of-Experts models locally.

Infrastructure and evaluation tools also evolved to address the needs of agentic systems. Platforms like Artificial Analysis launched Optima to help enterprises build custom performance evaluations, while Vals raised funding to expand its testing suites and cybersecurity benchmarks. Meanwhile, technical discussions highlighted various agent failure modes, such as the efficiency regressions caused by certain skill libraries and security concerns regarding leaked reasoning traces from proprietary application programming interfaces.

Read the original ↗