AI FRONTIER
Signals worth understanding before they become obvious.
Google releases RRSI, which lets an agent harness self-improve without retraining the model
Google Cloud AI Research researchers, with Stanford University and the University of Washington, released the RRSI framework, which repeatedly revises the harness around a model, including prompts, tool use, control flow, memory and context management, skill modules and sub-agents, without changing model weights.
- Across eight benchmarks, Terminal-Bench 2.1 rose from 74.2% to 80.2%, a 6.0 percentage point gain, and SWE-Bench Verified rose from 82.0% to 83.8%, a 1.8 percentage point gain.
- JobBench rose from 36.0% to 40.7%, a 4.7 percentage point gain; GDPval rose from 48.8% to 52.3%, a 3.5 percentage point gain; Frontier-Eng rose from 17.7% to 22.0%, a 4.3 percentage point gain.
- In a separate experiment using Gemini 3.5 Flash as the base model, Terminal-Bench 2.1 rose from 64.6% to 78.7%, a 14.1 percentage point gain, and SWE-Bench Verified rose from 76.8% to 79.0%, a 2.2 percentage point gain.
- RRSI is released as a research framework on GitHub, with code available under the Apache 2.0 license.
For teams building agents, RRSI shows that controlled iterative search over the harness, with model weights fixed, can deliver measurable benchmark gains and lower token cost.
Introducing GPT-6.1 Sol
OpenAI introduces GPT-6.1 Sol for coding, computer use, and professional work, reporting near-Astra intelligence at one-fifth of Astra's standard API input and output token prices.
- OpenAI reports GPT-6.1 Sol delivers near-Astra intelligence.
- Target uses are coding, computer use, and professional work.
- Standard API input and output token prices are one-fifth of Astra's.
For AI builders, near-Astra intelligence at one-fifth the token price lowers the cost per unit of inference on coding and computer-use workloads.
prism-ml releases Ternary-Bonsai-2-27B-gguf: a 27B ternary-weight model running from 5.95 GB
prism-ml published Ternary-Bonsai-2-27B-gguf on Hugging Face, quantizing the Qwen3.8-27B embeddings, attention projections, MLP projections and LM head to ternary weights (g128, {−1,0,+1} with FP16 group scaling) and shipping two GGUF packings, PTQ1_0 and PQ2_0, for llama.cpp on CUDA, Metal and CPU.
- Language model size is 5.95 GB (PTQ1_0) or 7.21 GB (PQ2_0), roughly 9.0x and 7.5x smaller than the ~54 GB FP16 baseline; the ideal ternary format is 1.72 bits/weight at 5.8 GB.
- The paper reports an 84.78 average across 14 thinking-mode benchmarks, 98.2% of the FP16 reference (86.32), versus 72.59 for IQ2_XXS (7.27 GB) and 85.18 for UD-Q4_K_XL (17.6 GB).
- The paper reports 95.83 on AIME26 and 90.07 on LiveCodeBench, where IQ2_XXS scores 57.5 and 56.4; math 96.57, coding 89.42, BFCL v3 tool calling 74.92.
- Context length is 262K tokens on a hybrid-attention backbone with about 75% linear attention; the vision tower is a separate Q8_0 mmproj pack (0.63 GB) loaded only for image input.
For teams deploying 27B-class reasoning on a single GPU or laptop, this shows end-to-end ternary quantization can hold near-FP16 reasoning and tool-calling scores at about 6 GB, but only with the dedicated kernels and forked runtime.
Towards Retrieving Interaction Spaces for Agentic Search: RISE builds a bounded space with BM25
The paper proposes RISE (Retrieving Interaction SpacE), which uses BM25 to construct a bounded subset of the corpus for an agent to explore and processes its documents during indexing for shell-style navigation. On BrowseComp-Plus, RISE matches the pure-shell DCI baseline at 78% accuracy with gpt-5.4-mini at roughly on...
- The paper reports that on BrowseComp-Plus, RISE reaches 78% accuracy with gpt-5.4-mini, matching the pure-shell DCI baseline at roughly one quarter of the per-query cost.
- The paper reports that at 1M documents, RISE-BM25 reaches 81% on gpt-5.4-mini.
- The paper reports that at the same 1M-document scale, DCI on gpt-5.4-nano degrades to 60% with 33 of 100 wall-clock failures.
- The paper argues that unbounded interaction does not scale: every broad shell command is a scan over the whole corpus, and latency degrades sharply as the corpus grows.
For teams building search agents, the result indicates retrieval should bound the interaction space rather than only select documents that fit the context window, which controls cost and latency as corpora grow.
Diffusion Controller unifies and simplifies AI image generation
Google Research introduces Diffusion Controller, a lightweight "steering damper" network that dynamically adjusts the denoising trajectory while the base model stays frozen, improving prompt alignment and image quality.
- Evaluated on a Stable Diffusion v1.4 backbone across supervised fine-tuning (SFT), reward-weighted loss (RWL), and PPO.
- In the SFT and RWL tracks, the gray-box Diffusion Controller outperformed LoRA in HPS-v2 win rates while manipulating significantly fewer internal model layers.
- The fully unlocked white-box version achieved a 90% win rate over the baseline model.
- A single inference-time guidance strength parameter allows dynamic adjustment of control constraint intensity.
For builders, this offers a lightweight path to controllable image generation and personalization without accessing closed-source model weights.
Google DeepMind introduces SynthID Bio for watermarking AI-generated proteins
Google DeepMind released SynthID Bio, a family of watermarking methods for synthetic biology that embeds a detectable signature in amino acid sequences and AlphaFold 3-predicted 3D structures while preserving biological function in laboratory testing.
- SynthID Bio fine-tunes a small part of AlphaFold 3's diffusion network, building watermarking into the model's weights so predicted 3D coordinates inherently carry a detectable signature.
- In wet-lab testing across three targets (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), watermarked designs matched the hit rate, binding affinity, and natural sequence diversity of unwatermarked versions.
- The work reports that SynthID Bio preserves AlphaFold 3 prediction accuracy while offering near-perfect detectability and holding up against digital noise or minor coordinate changes.
- Google DeepMind says it is publishing its methods paper, open-sourcing the code and in vitro data, and releasing the weights to the research community.
For builders of biological design tools, embedding watermarks directly in model weights and sequence generation offers an automated verification layer for DNA synthesis screening and database provenance, though robustness against deliberate...