01AgentsProduct4 sources agree
Glean has developed Waldo, a specialized agentic search model that works in concert with frontier models to deliver end-to-end agentic outcomes, reducing latency by 50% and token usage by 25% while maintaining quality. Waldo is trained using a combination of direct preference optimization and reinforcement learning, and is designed to determine how much reasoning a task requires based on its own execution. The model will be deployed to customers soon.
Provenance — who else covered this
02AgentsProduct4 sources agree
Glean introduces its third-generation AI Assistant with advanced personalization and agentic intelligence, and announces significant expansions to its Work AI platform, including new SDKs and MCP capabilities. The new Enterprise Graph underpins these innovations, enabling AI that truly understands the enterprise and its workflows. Glean Assistant now delivers complete outcomes tailored to each employee's way of working, without requiring advanced prompt engineering. The platform also features a new interactive workspace, Canvas, and allows employees to control their Assistant experience. Additionally, Glean is changing the way agents are built and deployed, making it easier for anyone to create and refine agents, and adding richer actions and MCP directory and host support.
Provenance — who else covered this
03AgentsProduct2 sources agree
Miles v0.1, a new open-source RL framework, has been announced, and multiple developments in agent evaluation, search benchmarking, and harnesses have been reported, including LangSmith Tuned Evaluators and Managed Deep Agents, indicating a shift towards robust rollouts, CI, observability, and environment plumbing. Search benchmarking for agents is maturing, with Artificial Analysis launching its Search Index, and LangChain introducing LangSmith Tuned Evaluators, which claim better performance at lower cost.
Provenance — who else covered this
04BusinessProduct3 sources agree
Glean's model routing helps control AI costs for organizations by selecting the most cost-effective model for each task, and its human feedback loop improves the routing system. The company has seen significant growth, reaching $300 million in annual recurring revenue, and is becoming a key player in the enterprise AI market. Glean's architecture includes a model called Waldo, which filters user queries and determines the best model to use, and the company is also seeing increased interest in open-weight models due to cost concerns.
Provenance — who else covered this
05ModelsProductsingle source
Qwen3.8-27B has become the top local model in Cline, with impressive benchmark results, while GLM-5.3 has been launched via API with significant gains in intelligence index and Elo rating, driven by stronger post-training techniques. The development of capable local models raises safety implications and highlights a shift towards useful, locally deployable models. GLM-5.3's gains suggest a shift in agentic capability scaling from parameter count to RL systems and environment quality.
Provenance — who else covered this
06SafetyProductsingle source
OpenAI paused some frontier RL training for two weeks to strengthen monitoring, isolation, and red-teaming, and is holding its largest planned frontier RL run, with a focus on hardening security and alignment controls. The slowdown mainly affects farther-out releases, not models already near ship.
Provenance — who else covered this
07AgentsProductsingle source
The Public AI Observatory is a new effort to measure real AI assistant usage, with 24,521 consented conversations and 52 models analyzed, while separate research highlights key findings on multi-agent teams and training variance, including the impact of task structure on communication topology and the emergence of specification gaming in agent collectives. The observatory aims to provide public-interest observability for AI usage patterns, independent of vendor reporting.
Provenance — who else covered this
08On-deviceProduct2 sources agree
Modular has open-sourced Mojo, positioning it as a portability layer across accelerators, while NVIDIA launched TensorRT Model Connect for direct model conversion and deployment, and Cursor published a retrospective on Git hosting at scale. Additionally, several companies announced advancements in on-device inference and datacenter accelerators, including DFlash 2 and Cerebras CS-4, highlighting the increasing importance of inference speed.
Provenance — who else covered this
09AgentsProduct3 sources agree
The DeepSeek Harness repository has achieved a record-breaking 20k stars in approximately one hour, surpassing previous records set by AutoGPT and Grok-1
Provenance — who else covered this
10CodingProductsingle source
Glean now allows users to select a specific AI model for each Assistant conversation, with options including GPT, Claude, and Gemini models, and also offers an Auto option that automatically selects the best model based on internal evaluations and live usage data. Admins can control which models are available for their users and set default models for their organization.
Provenance — who else covered this
11AgentsProductsingle source
Anthropic's Claude autonomously designed protein binders for 14 out of 15 targets, and Claude gained Gmail and Google Drive actions, while other AI developments and updates were also announced
Provenance — who else covered this
12ModelsInternals4 sources agree
DeepSeek V4 Pro is now available in Perplexity Computer, offering a cost-effective solution with a 62% cheaper cost-performance ratio compared to the next model, and it scored 0.359 at $0.75 per task on WANDR
Provenance — who else covered this
13SafetyProduct5 sources agree
AI models are experiencing issues with following instructions, a problem that may not be rapidly solved, similar to the challenge of hallucinations
Provenance — who else covered this
14CodingProduct2 sources agree
Anthropic's Claude Code has seen significant growth, now accounting for around 4% of GitHub code, and its usage is discussed in the context of AI engineering and semi-analysis work. The episode also touches on Memory Mania and its potential impact on users.
Provenance — who else covered this
15CodingProduct2 sources agree
Coding agents have improved the development process by reducing barriers between different teams, allowing for faster creation of working apps
Provenance — who else covered this
16AgentsProduct2 sources agree
Glean's harness and routing capabilities were benchmarked against Claude Cowork, showing a 4x cost advantage due to lower token volume and cheaper rates. Glean achieved this through model family routing, model tier routing, and better context handling, consuming fewer tokens per query. More results will be presented at Glean:GO!
Provenance — who else covered this
17CodingProduct2 sources agree
OJO streamlines the process of turning ideas into browser-accessible tools by handling the layer above coding agents, reducing the need for multiple handoffs and extensive development work
Provenance — who else covered this
18AI securityProduct2 sources agree
Zai released GLM 5.3, which was evaluated by Semgrep for its vulnerability detection capabilities, and the results show potential for open weight models in cyber security tasks
Provenance — who else covered this
19AgentsProduct2 sources agree
Zillow implemented Glean to unify its fragmented data landscape, enabling employees to quickly find information and deploy specialized agents across critical workflows, resulting in improved onboarding, customer focus, and engineering acceleration. Glean's integration with MCP servers and AI coding assistants like Claude Code also accelerated project initiation and raised coding standards.
Provenance — who else covered this
20BusinessBig picturesingle source
Booking.com implemented Glean's AI and search platform to improve information access and productivity, resulting in reduced video script creation time and faster IT ticket resolution. The company also integrated AI into its strategy and workflows, adopting Glean as its first company-wide AI platform
Provenance — who else covered this
21CodingProductsingle source
Glean has updated its Model Hub to allow configuration of available models and selection of models for workflows, with options for Glean Universal Model Key and Customer Key deployments. The update includes best practices for model selection and management.
Provenance — who else covered this
22RoboticsProductsingle source
A new paper presents a framework that unifies 3DGS captured assets using Gaussian Splats, potentially improving physics understanding in robots. This could lead to advancements in areas like table tennis-playing robots.
Provenance — who else covered this
23AgentsProductsingle source
Rox developed agents that automate research, outreach, deals, and renewals on top of a CRM backend, with successful implementations on MongoDB and Together
Provenance — who else covered this
24AI securityProduct3 sources agree
Semgrep's GLM 53 delivers Opus 48 level cybersecurity results at a lower cost, as analyzed by the company's colleagues
Provenance — who else covered this
25CodingProductsingle source
CC AI productivity agent in Gmail opens waitlist in Australia and New Zealand, and expands availability in the US and Canada, rolling out invitations to waitlisted users
Provenance — who else covered this
26BusinessBig picturesingle source
Memory prices have increased by 500% in the past 12 months, with 128GB DDR5 kits now costing $3,399, due to high demand from AI datacenter buildouts and limited supply. This price surge is affecting not only RAM but also other components like hard drives and SSDs.
Provenance — who else covered this
27AgentsProductsingle source
Perplexity's Computer now allows users to interact with it via email, enabling send, forward, or cc actions on any thread, with tasks running as normal sessions and maintaining an audit trail
Provenance — who else covered this