OpenAI has reduced prices for its GPT-5.6 model family. The Luna variant becomes 80% cheaper, while the Terra variant sees a 20% price cut. The Sol Fast inference service is now up to 2.5× faster. These updates are seen as a direct competitive move against Chinese AI models. Meanwhile, Grok 4.5 High remains on the Pareto frontier of cost-performance.
A user of the 'opencode go' subscription service received an HTTP 403 error when attempting to use DeepSeek V4 Flash. The error body is a RegionError stating that the latest version of the model is only available hosted in China and requires explicit opt-in. This indicates that DeepSeek has geo-restricted the V4 Flash model to Chinese IP addresses or accounts. The restriction affects non-China users even when using a third-party service that previously provided access. The situation may signal a change in DeepSeek's deployment policy for its latest models.
ThinkyMachines has announced Inkling-Small, an open omni-modal model capable of processing audio, text, and image inputs. The model is described as a frontier open Omni. It is available to run on Inference Endpoints, simplifying deployment for developers. This release expands open access to multimodal AI models beyond text and image to include audio.
Thinking Machines has released Inkling-Small, a new 276 billion parameter open-source model. UnslothAI announced that the model can now be run, making it accessible for inference. Inkling-Small is claimed to be the strongest open model for its size class.
Thinking Machines has released a new model called Inkling Small. The model uses NVFP4 precision and features 12 billion active parameters out of a total of 276 billion parameters. The source claims the model performs better than larger models, though no specific benchmarks or comparisons were provided.
On August 6, Together Compute will host a live walkthrough of major updates to its inference platform designed to simplify production deployment of open-weight models. The new features include safe rollout strategies (canary, blue-green, and automatic rollback), A/B and shadow testing on live traffic, and SLO-driven autoscaling. The platform also now supports deploying any model with approximately 4x faster warm starts. The session will be led by inference product manager Nikitha and developer relations lead Zain Has.