01On-deviceProduct2 sources agree
Unsloth has made GLM-5.3, the strongest open model to date, available for local runs, achieving 81% accuracy after significant model size reduction. The model can be run on a 256GB Mac or equivalent RAM/VRAM setups.
Provenance — who else covered this
02BusinessProductsingle source
Anthropic unexpectedly cut off most first-party capacity to Claude 3.x models with short notice, but the affected party has secured sufficient near-term capacity from other providers. Other models like Gemini 2.5 Pro and GPT 4.1 remain unaffected.
Provenance — who else covered this
03SafetyProductsingle source
Claude was tested on safety benchmarks for common misalignments and its best methods were evaluated on held-out benchmarks for generalization.
Provenance — who else covered this
04ModelsInternalssingle source
TencentHunyuan's Hy4 preview has achieved a significant improvement in the Code Arena, ranking #5 with 1633 points, compared to Hy3's #31 ranking. The score is based on an early AutoEval assessment, which uses a Reward Model trained on human preference data.
Provenance — who else covered this
05ModelsProduct2 sources agree
A configuration update for GLM-5.3-Flash aims to improve performance in certain agentic use cases, particularly for workflows where it previously underperformed Ox Alpha.
Provenance — who else covered this
06AgentsProduct2 sources agree
Grok Bot introduces a new harness strategy, involving daily automated checks by Tess, with Cody and Desi ensuring alignment with rules and design guidelines. The process automates verification of sites.
Provenance — who else covered this
07BusinessBig picturesingle source
Anthropic has rebranded CLIO, the research that inspired LangSmith Insights, as Anthropic Insights. LangSmith Insights was internally known as CLIO, which was the original name for the research.
Provenance — who else covered this
08AgentsProduct7 sources agree
Agents can reach websites in 6 different ways, with Google and Microsoft focusing on one method, and a demonstration of an agent buying something on a website is provided.
Provenance — who else covered this
09ResearchInternals3 sources agree
Qwen3.8-Flash-Next has difficulty tracking conversations over multiple turns, answering questions from previous turns at FP8.
Provenance — who else covered this
10AgentsProduct2 sources agree
AI models are being trained to perform complex tasks in realistic settings, including long-horizon agentic tasks.
Provenance — who else covered this