ragornot

News

The retrieval and AI landscape — RAG, LLMs, cost, efficiency, and the models behind them. Refreshed hourly. See the Benchmark tab for the empirical cost and quality numbers behind these technologies.

Refreshing feed…

LLMarXiv (cs.CL)

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through

CostarXiv (cs.CL)

Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures

RAGarXiv (cs.CL)

A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding

Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there still exist challenges for detecting out-of-scope (OOS) intents. (i) The traditional

LLMarXiv (cs.CL)

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). However, widely used critic-free RL methods rely on uniform credit assignment, broadcas

LLMarXiv (cs.CL)

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synthesize training data, including rational

LLMAWS Machine Learning Blog

Disaggregated prefill and decode for LLM inference on SageMaker HyperPod

In this post, we show how to implement DPD with vLLM on Amazon SageMaker HyperPod using the HyperPod Inference Operator.

LLMSimon Willison's Blog

llm-meta-ai 0.1

Release: llm-meta-ai 0.1 Let's LLM run prompts against the new muse-spark-1.1 model. Tags: llm , meta

LLMSimon Willison's Blog

llm 0.31.1

Release: llm 0.31.1 Fix for a bug with OpenAI Chat Completion endpoints where a tool call with empty arguments could result in a JSON error from some providers. #1521 This bug came up when I was testing llm-meta-ai . Tag

RAGAWS Machine Learning Blog

Powering scientific discovery: BYOKG and GraphRAG for intelligent pharmaceutical research

In this post, we explore how Graph-based Retrieval Augmented Generation (GraphRAG) is transforming scientific research by combining graph databases with generative AI. With this approach, you can accelerate discovery pro

CostHacker News

From Words to Watts: Benchmarking the Energy Costs of LLM Inference (2023)

LLMHugging Face Blog

Native-speed vLLM transformers modeling backend

RAGQdrant Blog

Qdrant Beats Elastic’s DiskBBQ at 2x Throughput, Half the Latency, and 1/3 the Compute

TL;DR Elastic recently published a benchmark claiming that their proprietary, disk-based index (dubbed “DiskBBQ”) delivers up to 7x higher throughput than Qdrant when deployed on nodes with network-attached storage. Elas

LLMSimon Willison's Blog

github-code Web Component

Tool: github-code Web Component An experimental Web Component built using GPT-5.5 and the following prompt : let's build a Web Component for embedding code from GitHub <github-code href="https://github.com/simonw/sqli

RAGHugging Face Blog

Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot

LLMSimon Willison's Blog

llm-coding-agent 0.1a0

Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would look like built on it. I started a new Pyt

LLMHacker News

Show HN: fenic – LLMs as dataframe operators, query meaning and structure

LLMAhead of AI

Using Local Coding Agents

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions

RAGHacker News

Show HN: RAG Vector DB Cost Calculator

LLMHugging Face Blog

Run a vLLM Server on HF Jobs in One Command

LLMHacker News

Show HN: Sipp – Run small local LLMs in browser 3x faster

LLMOpenAI

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

RAGHugging Face Blog

Experimenting with the proposed Cross-Origin Storage API in Transformers.js

CostHacker News

Quantifying LLM Cost Savings from Cache-Aware Inference Routing

EnvironmentGoogle AI Blog

Our new community investments in Virginia support local jobs and expand energy affordability.

We’re helping build the state’s next-generation workforce and investing in energy programs.

LLMAhead of AI

LLM Research Papers: The 2026 List (January to May)

A curated roundup of notable LLM research papers that came out this year

LLMHugging Face Blog

Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic

RAGOpenAI

How Virgin Atlantic ships faster with Codex

How Virgin Atlantic used Codex to ship its revamped mobile app on a fixed holiday travel deadline, reaching near-total unit test coverage and zero P1 defects.

RAGWeaviate Blog

Build a Coding Assistant with Weaviate MCP: RAG over Code & Docs

Use Weaviate's built-in MCP server to give Claude Code, Cursor, and VS Code hybrid search over your codebase and docs. No glue code.

EnvironmentGoogle AI Blog

We’re announcing new community investments in Missouri.

We’re helping build the state’s next-generation workforce and investing in energy programs.

AIVentureBeat AI

Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

For a quarter century, the Google search box has been one of the most recognizable interfaces in computing: a thin white rectangle, a blinking cursor, a few typed words, and a list of blue links. On Tuesday, Google will

RAGQdrant Blog

How GoPerfect Built an Agentic Recruiting Workforce with Qdrant Cloud

GoPerfect mission is to use an AI recruiting workforce that replaces the manual, low-leverage parts of recruiting. Instead, an agent decomposes recruiter intent and runs the work end to end to find top talent. Their agen

CostAhead of AI

Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention

From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs

EnvironmentDeepMind Blog

Strengthening Singapore’s AI Future: A New National Partnership

Google DeepMind and Singapore partner to apply frontier AI to address complex challenges across health, education, and sustainability and more.

RAGOpenAI

Parloa builds service agents customers want to talk to

Parloa leverages OpenAI models to power scalable, voice-driven AI customer service agents, enabling enterprises to design, simulate, and deploy reliable, real-time interactions.

RAGWeaviate Blog

Your LLM Is Only as Good as What It Retrieves

A Researcher's Perspective on Retrieval Quality in RAG Systems

RAGQdrant Blog

How Data Graphs Built a True Hybrid Graph RAG Platform

Data Graphs is a UK-based platform company that provides a knowledge graph-as-a-service polystore, built on a proprietary, high-performance graph database engine. Co-founded by Paul Wilton over 10 years ago as a consulta

LLMAhead of AI

My Workflow for Understanding LLM Architectures

A learning-oriented workflow for understanding new open-weight model releases

LLMOpenAI

AI fundamentals

Learn what AI is, how it works, and how tools like ChatGPT use large language models. A clear, beginner-friendly guide to understanding artificial intelligence.

RAGOpenAI

CyberAgent moves faster with ChatGPT Enterprise and Codex

CyberAgent uses ChatGPT Enterprise and Codex to securely scale AI adoption, improve quality, and accelerate decisions across advertising, media, and gaming.

LLMAhead of AI

Components of A Coding Agent

How coding agents use tools, memory, and repo context to make LLMs work better in practice

RAGWeaviate Blog

Multimodal Embeddings and RAG: A Practical Guide

Multimodal embeddings allow AI systems to search and reason across text, images, audio, and video in their native formats. This blog covers the key intuitions behind how this all works and walks through three practical i

RAGWeaviate Blog

Your Code is Your Schema: Weaviate Managed C# Client

Use semantic search and RAG in C# with the Weaviate Managed .NET client — attribute-driven schema, type-safe queries, and safe migrations, all in idiomatic .NET.

RAGQdrant Blog

Qdrant Skills for AI Agents

The standard RAG tutorial teaches a simple pattern: embed your documents, store them in a vector database, retrieve the top K, and feed them to the LLM. The vector engine is passive infrastructure. Put vectors in, get ne

RAGQdrant Blog

Master Multi-Vector Search With Qdrant

Most vector search tutorials stop at single-vector embeddings: one document, one vector, one similarity score. That works for demos. It falls apart when your retrieval pipeline needs to capture fine-grained token-level i

RAGWeaviate Blog

Building A Legal RAG App in 36 Hours

Learn how we built a production-ready, end-to-end RAG application in just 36 hours using the Query Agent and the new Weaviate Agent Skills library.

AIThe Gradient

After Orthogonality: Virtue-Ethical Agency and AI Alignment

Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but be

CostVentureBeat AI

Railway secures $100 million to challenge AWS with AI-native cloud infrastructure

Railway , a San Francisco-based cloud platform that has quietly amassed two million developers without spending a dollar on marketing, announced Thursday that it raised $100 million in a Series B funding round, as surgin

CostVentureBeat AI

Claude Code costs up to $200 a month. Goose does the same thing for free.

The artificial intelligence coding revolution comes with a catch: it's expensive. Claude Code , Anthropic's terminal-based AI agent that can write, debug, and deploy code autonomously, has captured the imaginat

AIVentureBeat AI

Listen Labs raises $69M after viral billboard hiring stunt to scale AI customer interviews

Alfred Wahlforss was running out of options. His startup, Listen Labs , needed to hire over 100 engineers, but competing against Mark Zuckerberg's $100 million offers seemed impossible. So he spent $5,000 — a fifth

AIVentureBeat AI

Salesforce rolls out new Slackbot AI agent as it battles Microsoft and Google in workplace AI

Salesforce on Tuesday launched an entirely rebuilt version of Slackbot , the company's workplace assistant, transforming it from a simple notification tool into what executives describe as a fully powered AI agent c

LLMDeepMind Blog

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.

EnvironmentDeepMind Blog

Google DeepMind supports U.S. Department of Energy on Genesis: a national mission to accelerate innovation and scientific discovery

Google DeepMind and the DOE partner on Genesis, a new effort to accelerate science with AI.

RAGDeepMind Blog

How AI is giving Northern Ireland teachers time back

A six-month long pilot program with the Northern Ireland Education Authority’s C2k initiative found that integrating Gemini and other generative AI tools saved participating teachers an average of 10 hours per week.

LLMDeepMind Blog

T5Gemma: A new collection of encoder-decoder Gemma models

Introducing T5Gemma, a new collection of encoder-decoder LLMs.

AIThe Gradient

AGI Is Not Multimodal

"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of generative AI models have convinc

AIThe Gradient

Shape, Symmetries, and Structure: The Changing Role of Mathematics in Machine Learning Research

What is the Role of Mathematics in Modern Machine Learning? The past decade has witnessed a shift in how progress is made in machine learning. Research involving carefully designed and mathematically principled architect

LLMThe Gradient

What's Missing From LLM Chatbots: A Sense of Purpose

LLM-based chatbots’ capabilities have been advancing every month. These improvements are mostly measured by benchmarks like MMLU, HumanEval, and MATH (e.g. sonnet 3.5, gpt-4o). However, as these measures get more

AIThe Gradient

We Need Positive Visions for AI Grounded in Wellbeing

Introduction Imagine yourself a decade ago, jumping directly into the present shock of conversing naturally with an encyclopedic AI that crafts images, writes code, and debates philosophy. Won’t this technology al

ragornot aggregates headlines and links to original sources. Full articles live on the publisher’s site.