AI Agent Memory Is a Dose, Not a Switch
IBM tested agent memory across eight models. More memory made weaker models worse. How to calibrate the dose, and why prompt caching decides the bill.
IBM tested agent memory across eight models. More memory made weaker models worse. How to calibrate the dose, and why prompt caching decides the bill.
My eval harness leaked model names into the judge prompt for six weeks. What two recent papers say about judge bias and calibration, and the…
New research shows LLM API testing misses how chatbots actually behave. Here's what the data says about the gap and how to fix your eval…
A new benchmark shows LLMs drop 0.3 to 5.9% accuracy on grade-school math when names and places go non-Western. Same arithmetic, different answers.
A paper trains a tiny probe on a model's own hidden states to catch hallucinations at inference time, no judge model required. Here's why that…
Most reasoning LLM failures aren't hallucinations, they're silently skipped steps. Here's what to measure instead of end-to-end answer accuracy.