TMtechmemeGPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8% (Greg Kamradt/ARC Prize)industry10 days ago▲ 24Read story→
TMtechmemeGoogle rolls out its September Android Drop, with remembered items in Find Hub, Guided vision in Gemini Live, Motion Assist to reduce motion sickness, and more (Ryan Whitwam/Ars Technica)industry11 days ago▲ 24Read story→
arXivarxivBeyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluationarxivcs.CLcs.LG12 days ago▲ 43Read story→
arXivarxivClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialoguesarxivcs.CLpublisher:arxiv12 days ago▲ 43Read story→
arXivarxivLongPIBench: A Long-Context Benchmark for Prompt Injectionarxivcs.CRcs.AI16 days ago▲ 43Read story→
arXivarxivPersona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Auditarxivcs.SEcs.AI17 days ago▲ 43Read story→