An auto-updating feed of the newest AI research, gathered from the top sources and refreshed on a schedule. Freshly added papers are marked 🆕.
250 papers from the last 30 days · Last updated 2026-07-30 12:09 UTC
Where the papers below come from, plus the major labs and venues worth watching.
| Source | What it is |
|---|---|
| arXiv (cs.AI / cs.LG / cs.CL / cs.CV) | The primary open preprint server — where nearly all AI research lands first. (Auto-tracked) |
| Hugging Face Papers | Community-curated daily & trending papers. Official successor to Papers with Code. (Auto-tracked) |
| OpenReview | Open peer review for ICLR, NeurIPS, and other top venues. |
| Semantic Scholar | AI-powered scholarly search with citation graph and a free API. |
| OpenAlex | Fully open index of the world's scholarly papers; free API. |
| Google DeepMind — Research | Publications from one of the leading industrial AI labs. (Watch — no reliable public feed yet) |
| OpenAI — Research | Research index and technical reports from OpenAI. (Auto-tracked via RSS) |
| Anthropic — Research | Papers on safety, interpretability, and frontier models. (Auto-tracked via scraping — best-effort) |
| Meta AI (FAIR) — Research | Fundamental AI Research publications from Meta. (Auto-tracked via scraping — best-effort) |
| Google Research | Broad publications across ML, systems, and applications. (Auto-tracked via RSS) |
| Microsoft Research | Extensive AI and systems research output. |
| Papers with Code (archived) | Retired by Meta in July 2025; historical SOTA/benchmark data is frozen here. |
Why these are worth your time — picked from the top of the feed.
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone — HiFi-UMI enables direct deployment of manipulation policies trained solely on high-fidelity UMI data, eliminating the need for costly real-robot post-training anchors.
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM — TurboVLA radically improves VLA efficiency by bypassing LLMs, achieving high performance with minimal parameters and VRAM usage, ideal for real-time applications.
- CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents — CodeNib achieves up to 25.4x faster context updates and 87% fewer tokens in navigation, drastically improving coding agents' efficiency with evolving repositories.
- CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization — CoRT enables refined credit assignment in language model outputs by using counterfactual replay, significantly improving reward distribution across individual tokens without altering the overall reward structure.
- HumanCLAW: Can Vision-Language Models Act Through a Body? — HumanCLAW introduces a novel framework to isolate decision-making in vision-language models from execution errors, revealing significant performance limitations in state-of-the-art VLMs.
- ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition — ReDesign's agentic framework significantly improves editability from raster images by integrating graceful verification, achieving superior results over existing methods in both fidelity and functionality.
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space — DecoEvo introduces a novel decoupled approach for evolving solver and rubric-generator skills, allowing for more effective optimization in open-ended text tasks by addressing overlooked metrics.
- CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition — CLBench-V offers a comprehensive evaluation framework for multimodal context learning, providing critical insights into grounding and knowledge acquisition across diverse real-world tasks.
- Pass the Baton: Trajectory-Relayed On-Policy Distillation — Relay-OPD significantly outperforms standard on-policy distillation by addressing prefix failure through strategic teacher intervention, enhancing performance across mathematical reasoning benchmarks.
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents — CAST leverages game solver state value changes as turn-level signals, significantly enhancing reinforcement learning performance in long-horizon games across multiple domains.
The newest papers (80 shown), most recent first — each with a direct link and a short description.
2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲109
Hengyi Xie, Chenfei Yao, Xianjin Wu et al.
Vision-language-action (VLA) models commonly adopt an LLM-centric
🔗 https://arxiv.org/abs/2607.27205
2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲59
Siyao Li, Jiawei Gu, Shuai Liu et al.
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed…
🔗 https://arxiv.org/abs/2607.27180
2026-07-29 · ✨ Other AI Research · Hugging Face · ▲15
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu et al.
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one ta…
🔗 https://huggingface.co/papers/2607.26784
2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲10
Zijun Lin, Zeqing Wang, Cheston Tan et al.
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activat…
🔗 https://arxiv.org/abs/2607.26754
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding 🆕
2026-07-29 · 🤖 Artificial Intelligence (General) · Hugging Face, arXiv · ▲6
Jingbo Zhou, Yusai Zhao, Qi Bao et al.
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a bench…
🔗 https://arxiv.org/abs/2607.27155
2026-07-29 · 🤖 Artificial Intelligence (General) · Hugging Face, arXiv · ▲5
Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generat…
🔗 https://arxiv.org/abs/2607.27191
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Julien Benchek, Austin Bennett, Jasmin Kern et al.
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set com…
🔗 https://arxiv.org/abs/2607.27189
The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making 🆕
2026-07-29 · 🧑💻 Human-AI Interaction · arXiv
Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos et al.
Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analy…
🔗 https://arxiv.org/abs/2607.27179
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Peter Tisnikar, Maja Swieczkowska, Benteng Ma et al.
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute…
🔗 https://arxiv.org/abs/2607.27177
2026-07-29 · 🔎 Information Retrieval & RAG · arXiv
Ji Xin, Xiao Xiao, Ishan Bhatt et al.
Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the di…
🔗 https://arxiv.org/abs/2607.27172
2026-07-29 · 👁️ Computer Vision · arXiv
Roshan Kenia, Stephanie L McNamara, William Lotter
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning ana…
🔗 https://arxiv.org/abs/2607.27154
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark 🆕
2026-07-29 · 🧠 Machine Learning · arXiv
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage gua…
🔗 https://arxiv.org/abs/2607.27143
2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Zuojin Tang, Feifan Luo, Haoyun Liu et al.
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations w…
🔗 https://arxiv.org/abs/2607.27138
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduc…
🔗 https://arxiv.org/abs/2607.27134
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Yiping Song, Jiaoyan Chen, Renate Schmidt et al.
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we…
🔗 https://arxiv.org/abs/2607.27130
2026-07-29 · 🔊 Audio & Speech · arXiv
Weijie Wu, Junbo Li, Lin Li et al.
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to d…
🔗 https://arxiv.org/abs/2607.27109
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context 🆕
2026-07-29 · 👁️ Computer Vision · arXiv
Zihan Deng, Chuanzhi Xu, Huiqi Liang et al.
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs…
🔗 https://arxiv.org/abs/2607.27084
2026-07-29 · 🧠 Machine Learning · arXiv
Yicheng Feng, Yan Zhang, Yan Cheng et al.
As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and priv…
🔗 https://arxiv.org/abs/2607.27083
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Yongjian Guo, Wanlun Ma, Lingyu Shen et al.
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values…
🔗 https://arxiv.org/abs/2607.27081
2026-07-29 · ✨ Other AI Research · arXiv
Xuanze Chen, Xukang Xie, Wentao Fu et al.
Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent ben…
🔗 https://arxiv.org/abs/2607.27080
2026-07-29 · 🧠 Machine Learning · arXiv
Vaneet Aggarwal
We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite
🔗 https://arxiv.org/abs/2607.27073
2026-07-29 · 👁️ Computer Vision · arXiv
Feixiang Liu, Qiang Qiu, Lanbo Sun et al.
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decisio…
🔗 https://arxiv.org/abs/2607.27069
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence 🆕
2026-07-29 · 👁️ Computer Vision · arXiv
Chuanzhi Xu, Zihan Deng, Huiqi Liang et al.
Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image…
🔗 https://arxiv.org/abs/2607.27066
2026-07-29 · 👁️ Computer Vision · arXiv
Paul Julius Kühn, Saptarshi Neil Sinha, Tiago Kleist et al.
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale anno…
🔗 https://arxiv.org/abs/2607.27065
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data 🆕
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Lingyang Zeng, Guangze Chen, Kaichen Yu et al.
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However,…
🔗 https://arxiv.org/abs/2607.27056
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation 🆕
2026-07-29 · 🧠 Machine Learning · arXiv
Fengming Yu, Haiwei Pan, Kejia Zhang et al.
Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences…
🔗 https://arxiv.org/abs/2607.27054
2026-07-29 · 🧠 Machine Learning · arXiv
Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet et al.
Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requi…
🔗 https://arxiv.org/abs/2607.27023
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception 🆕
2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Gabe Everett, Brice Gunter, Ryan Vander Stelt et al.
Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations t…
🔗 https://arxiv.org/abs/2607.26985
2026-07-29 · 👁️ Computer Vision · arXiv
Duzhen Zhang, Yahan Yu, Qiaoyi Su et al.
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving…
🔗 https://arxiv.org/abs/2607.26947
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand et al.
Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-gra…
🔗 https://arxiv.org/abs/2607.26946
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation 🆕
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Vishisht Choudhary, Lukas Schmidt, Anne Zoë Kenntner et al.
Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection f…
🔗 https://arxiv.org/abs/2607.26935
2026-07-29 · ✨ Other AI Research · arXiv
Hongliang Zhang, Zhongyuan Yu, Guijuan Wang et al.
Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthi…
🔗 https://arxiv.org/abs/2607.26933
2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Zhe Liu, Quan Lu, Zhaohui Du et al.
Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inad…
🔗 https://arxiv.org/abs/2607.26914
2026-07-29 · 🧠 Machine Learning · arXiv
Brandon Gower-Winter, Georg Krempl
In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. This phenomena is known as Outcome Performativity. This paper formalises an approach for detecting Outcome Performativity using predi…
🔗 https://arxiv.org/abs/2607.26908
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence 🆕
2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Jia Luo
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource f…
🔗 https://arxiv.org/abs/2607.26903
2026-07-29 · 🧑💻 Human-AI Interaction · arXiv
Mengchen Dong, Hiromu Yakura
Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges…
🔗 https://arxiv.org/abs/2607.26899
2026-07-29 · 👁️ Computer Vision · arXiv
Siddharth Vohra
When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and derm…
🔗 https://arxiv.org/abs/2607.26886
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents 🆕
2026-07-29 · 📈 ML Theory & Statistics · arXiv
Amirmohammad Farzaneh, Osvaldo Simeone
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remainin…
🔗 https://arxiv.org/abs/2607.26865
2026-07-29 · 🧠 Machine Learning · arXiv
Junoh Park, Junseo Hwang, Wonguk Cho et al.
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage…
🔗 https://arxiv.org/abs/2607.26862
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Ruikang Zhang, Shuo Wang, Qi Su
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have prima…
🔗 https://arxiv.org/abs/2607.26853
2026-07-29 · 🧠 Machine Learning · arXiv
Yansen Zhang, Yilu Liu, Tianyu Liu et al.
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cau…
🔗 https://arxiv.org/abs/2607.26828
2026-07-29 · 🧑💻 Human-AI Interaction · arXiv
Janin Koch, Xiaohan Liao, Géry Casiez
Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output translate into more value for the designer. We argue, however, that this framing leaves out something important about how d…
🔗 https://arxiv.org/abs/2607.26827
2026-07-29 · ✨ Other AI Research · arXiv
Wenhao Yang, Runzhi He, Minghui Zhou
Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents…
🔗 https://arxiv.org/abs/2607.26819
2026-07-29 · 🧠 Machine Learning · arXiv
Zhaoyang Ma, Zhihao Wu, Xin Gao et al.
Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. Existing…
🔗 https://arxiv.org/abs/2607.26801
2026-07-29 · 🧠 Machine Learning · arXiv
Perry Dong, Ron Polonsky, Dorsa Sadigh et al.
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional w…
🔗 https://arxiv.org/abs/2607.27203
2026-07-29 · 🧠 Machine Learning · arXiv
Shady E. Ahmed, Panos Stinis
We present a novel approach to regression tasks using classification which is motivated by the mechanism used by fruitflies to sense their environment. Specifically, we formulate a general framework for learning nonlinear input-output relationships by replacing complex global su…
🔗 https://arxiv.org/abs/2607.27196
2026-07-29 · 🧠 Machine Learning · arXiv
Lennon J. Shikhman, Michael Galarnyk, Aadi Dash et al.
Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only…
🔗 https://arxiv.org/abs/2607.27188
When Do Learned Diffusion Proposals Help Constraint Solving? A Controlled Study on Continuous Algebraic Systems 🆕
2026-07-29 · 🧠 Machine Learning · arXiv
Quang Bui, Sparsh Roy, Akash Gundimeda et al.
Solving a continuous algebraic constraint system requires two decisions: which values satisfy the constraints, and which structural augmentation renders an unsolvable system solvable. Classical solvers answer the first well and the second only by enumeration. On that discrete de…
🔗 https://arxiv.org/abs/2607.27169
2026-07-29 · 🧠 Machine Learning · arXiv
Francesco Pinto, Luca Lanzilao, Paco Lopez Dekker et al.
Accurate intraday forecasts of offshore wind are becoming increasingly important for power system operation and the integration of growing shares of offshore wind energy. Operational forecasts rely predominantly on numerical weather prediction (NWP), which is not optimized for l…
🔗 https://arxiv.org/abs/2607.27152
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis 🆕
2026-07-29 · ✨ Other AI Research · arXiv
Yihao Chen, Shi Chang, Khaled Chawa et al.
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on Prog…
🔗 https://arxiv.org/abs/2607.27146
Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devices 🆕
2026-07-29 · ✨ Other AI Research · arXiv
Harvey Samuel George Johnson, Sendy Phang
This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic. With the target workload of branch prediction (BP) for multistage pipelined central pro-cessing unit (CPU) cores. For thi…
🔗 https://arxiv.org/abs/2607.27140
2026-07-29 · 🧠 Machine Learning · arXiv
Zuyuan Zhang, Yongshan Chen, Mahdi Imani et al.
An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision pro…
🔗 https://arxiv.org/abs/2607.27132
2026-07-29 · 🧠 Machine Learning · arXiv
Kaifeng Zhang, Kai Ming Ting
Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves…
🔗 https://arxiv.org/abs/2607.27126
2026-07-29 · 🧠 Machine Learning · arXiv
Filipa Lino, Bárbara Tavares, Carlos Santiago et al.
Emergency Departments (EDs) are critical access points in healthcare systems, yet they face persistent pressure from unpredictable patient demand, seasonal surges, and non-urgent visits. Effective ED planning requires forecasts at multiple decision-making levels: hospitals need…
🔗 https://arxiv.org/abs/2607.27106
Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approach 🆕
2026-07-29 · ✨ Other AI Research · arXiv
Andrew Flynn, Cian McCafferty, Klaus Lehnertz et al.
Most existing seizure detection algorithms require extensive pre-processing of the data and rely on heuristic or currently unexplainable machine learning approaches. These approaches often struggle with balancing detection sensitivity and specificity in the presence of variable…
🔗 https://arxiv.org/abs/2607.27105
2026-07-29 · 🧠 Machine Learning · arXiv
Aleksandr Berdnikov, Yevgeny Liokumovich
We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. We find that most of the considered open-source models do have such a representation, and it often even surfaces to the top principal compone…
🔗 https://arxiv.org/abs/2607.27092
2026-07-29 · ✨ Other AI Research · arXiv
Peter Li, Prashant Pandey
Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant sub…
🔗 https://arxiv.org/abs/2607.27090
2026-07-29 · ✨ Other AI Research · arXiv
Hung Mai, Hai Nguyen, Luong Doan et al.
In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and scalar value-certified sampling. Our main result is a field-code compiler: any communicated transport field approximating an…
🔗 https://arxiv.org/abs/2607.27078
2026-07-29 · 🧠 Machine Learning · arXiv
Nicolas Béreux, Aurélien Decelle, Cyril Furtlehner et al.
Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the c…
🔗 https://arxiv.org/abs/2607.27077
Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework 🆕
2026-07-29 · 🧠 Machine Learning · arXiv
Kindeep K. Dhatt, Tengyue Wu, Hanbang Hua et al.
Continuous cuffless blood pressure (BP) monitoring remains challenging due to motion artifacts, physiological variability, and the limited robustness of conventional pulse transit time (PTT) models under dynamic conditions. Many prior approaches rely on multi-second windows to s…
🔗 https://arxiv.org/abs/2607.27076
2026-07-29 · 📈 ML Theory & Statistics · arXiv
Joachim Bona-Pellissier, Giacomo Meanti, Matteo Santacesaria et al.
Physics-informed machine learning incorporates physical principles --often expressed via differential operators-- into data-driven models. While physics-informed neural networks (PINNs) dominate empirical applications, the complexity of neural network architectures and optimizat…
🔗 https://arxiv.org/abs/2607.27062
2026-07-29 · ✨ Other AI Research · arXiv
Jiale Chen, Torsten Hoefler, Dan Alistarh
Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. They process the entries in a fixed order, one at a time, propagating each rounding error to the entries not yet processed through a…
🔗 https://arxiv.org/abs/2607.27042
2026-07-29 · 👁️ Computer Vision · arXiv
Taiye Chen, Qi Zhang, Yisen Wang
Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomeno…
🔗 https://arxiv.org/abs/2607.27036
2026-07-29 · 🧠 Machine Learning · arXiv
Bum Jun Kim
Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. Moreover, prior work has not considered actual deployment conditions, wh…
🔗 https://arxiv.org/abs/2607.27031
2026-07-29 · ✨ Other AI Research · arXiv
Petr Simecek, Elnaz Babayeva, Jiri Balhar et al.
LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark…
🔗 https://arxiv.org/abs/2607.27030
2026-07-29 · 🧠 Machine Learning · arXiv
James Chapman
Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-a…
🔗 https://arxiv.org/abs/2607.27027
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations 🆕
2026-07-29 · 🧠 Machine Learning · arXiv
Kaizhen Tan, Xin Xu, Siru Tao et al.
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEW…
🔗 https://arxiv.org/abs/2607.27017
2026-07-29 · ✨ Other AI Research · arXiv
Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina
Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare.…
🔗 https://arxiv.org/abs/2607.27000
2026-07-29 · ✨ Other AI Research · arXiv
Ruoyu Wang, Heng Zhao, Renjie Wu et al.
Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. H…
🔗 https://arxiv.org/abs/2607.26998
2026-07-29 · 🧠 Machine Learning · arXiv
Peter Lorenz, Anjith George, Sébastien Marcel
Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained vision models rather than training task-specific ar…
🔗 https://arxiv.org/abs/2607.26993
2026-07-29 · ✨ Other AI Research · arXiv
Franz Nowak, Ryan Cotterell, Reda Boumasmoud
What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains…
🔗 https://arxiv.org/abs/2607.26988
2026-07-29 · ✨ Other AI Research · arXiv
Michelangelo Domina, Michele Ceriotti
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc.,…
🔗 https://arxiv.org/abs/2607.26984
2026-07-29 · 📈 ML Theory & Statistics · arXiv
Yuheng Ma, Qiang Sun
We study feature bagging through the lens of algorithmic stability. Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner. We introduce feature instability (FI), the feature-axis…
🔗 https://arxiv.org/abs/2607.26964
2026-07-29 · 📈 ML Theory & Statistics · arXiv
Shuo-Chieh Huang, Chien-Ming Chi, Jau-er Chen
Minimax-optimal rates for multivariate distribution estimation are known to suffer from the curse of dimensionality. We propose a sparse Bayesian network approach in which each conditional probability is estimated using sparsity-aware conditional mean methods. The resulting esti…
🔗 https://arxiv.org/abs/2607.26955
2026-07-29 · 🧠 Machine Learning · arXiv
Alexandr Udeneev, Petr Babkin, Oleg Bakhteev
Ensembles are a standard way to improve the performance and robustness of deep neural networks, but their effectiveness crucially depends on both the quality and the diversity of individual models. Most neural architecture search (NAS) methods are computationally expensive. Exte…
🔗 https://arxiv.org/abs/2607.26940
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Hao Fei, Yiran Zhao
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, an…
🔗 https://arxiv.org/abs/2607.27201
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Ben Glickenhaus, Katherine Thai, Jenna Russell et al.
We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, P…
🔗 https://arxiv.org/abs/2607.27183
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search 🆕
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Raphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior et al.
State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first…
🔗 https://arxiv.org/abs/2607.27178
SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch 🆕
2026-07-29 · ✨ Other AI Research · arXiv
Yihao Chen, Shi Chang, Feng Lin et al.
LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an exec…
🔗 https://arxiv.org/abs/2607.27167
2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Jiayuan Di, Haoyi Yang, Yufei Luo et al.
Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce…
🔗 https://arxiv.org/abs/2607.27022
- 🗣️ NLP & Large Language Models (29)
- 🧠 Machine Learning (38)
- 🤖 Artificial Intelligence (General) (11)
- 👁️ Computer Vision (37)
- 🦾 Robotics & Embodied AI (25)
- 🧬 Neural & Evolutionary Computing (3)
- 🤝 Multi-Agent Systems (11)
- 🔎 Information Retrieval & RAG (25)
- 📈 ML Theory & Statistics (9)
- 🔊 Audio & Speech (1)
- 🧑💻 Human-AI Interaction (5)
- 🏛️ From the Labs (blogs & reports) (3)
- ✨ Other AI Research (53)
- APEX-Accounting 🆕
Julien Benchek, Austin Bennett, Jasmin Kern et al. · 2026-07-29 · arXiv - From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs 🆕
Ruikang Zhang, Shuo Wang, Qi Su · 2026-07-29 · arXiv - Mental World Modeling 🆕
Hao Fei, Yiran Zhao · 2026-07-29 · arXiv - Pangram 4 Technical Report 🆕
Ben Glickenhaus, Katherine Thai, Jenna Russell et al. · 2026-07-29 · arXiv - DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search 🆕
Raphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior et al. · 2026-07-29 · arXiv - Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making 🆕
Jiayuan Di, Haoyi Yang, Yufei Luo et al. · 2026-07-29 · arXiv - OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment 🆕
Seonglae Cho, Adriano Koshiyama · 2026-07-29 · arXiv - TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning 🆕
Jinhu Qi, Wentao Zhang, Siu Man Ng et al. · 2026-07-29 · arXiv - Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation 🆕
Weijie Feng, Hongchuang Wang, Binbin Liu et al. · 2026-07-29 · arXiv - Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? 🆕
Arnav Hiray, Agam Shah, Caleb Lu et al. · 2026-07-29 · arXiv - Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States 🆕
Weiyi Kong, Zhuoran Li · 2026-07-29 · arXiv - Latent-IM: Latent Interaction Management for Speech LLMs 🆕
Adar Avsian, Atahan Dokme, Tony Woo et al. · 2026-07-29 · arXiv - +17 more in this topic — see
data/papers.json
- Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark 🆕
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal · 2026-07-29 · arXiv - Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents 🆕
Yicheng Feng, Yan Zhang, Yan Cheng et al. · 2026-07-29 · arXiv - Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise 🆕
Vaneet Aggarwal · 2026-07-29 · arXiv - CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation 🆕
Fengming Yu, Haiwei Pan, Kejia Zhang et al. · 2026-07-29 · arXiv - BayesAME: Bayesian Active Model Evaluation 🆕
Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet et al. · 2026-07-29 · arXiv - Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing 🆕
Brandon Gower-Winter, Georg Krempl · 2026-07-29 · arXiv - ReCo: Reweighting GRPO Against Distributional Concentration 🆕
Junoh Park, Junseo Hwang, Wonguk Cho et al. · 2026-07-29 · arXiv - Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility 🆕
Yansen Zhang, Yilu Liu, Tianyu Liu et al. · 2026-07-29 · arXiv - FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning 🆕
Zhaoyang Ma, Zhihao Wu, Xin Gao et al. · 2026-07-29 · arXiv - Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? 🆕
Perry Dong, Ron Polonsky, Dorsa Sadigh et al. · 2026-07-29 · arXiv - From Classification to Regression: Using a Fruitfly to Solve Equations 🆕
Shady E. Ahmed, Panos Stinis · 2026-07-29 · arXiv - Inverse Learning of Latent Risk-Neutral Densities from Irregular Option Quotes 🆕
Lennon J. Shikhman, Michael Galarnyk, Aadi Dash et al. · 2026-07-29 · arXiv - +26 more in this topic — see
data/papers.json
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding 🆕
Jingbo Zhou, Yusai Zhao, Qi Bao et al. · 2026-07-29 · Hugging Face, arXiv · ▲6 - Can AI agents conduct open-ended AI research? Early evidence from two case studies 🆕
Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al. · 2026-07-29 · Hugging Face, arXiv · ▲5 - Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork 🆕
Peter Tisnikar, Maja Swieczkowska, Benteng Ma et al. · 2026-07-29 · arXiv - Linguistic Monoculture in LLM-Assisted Language Use 🆕
Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen · 2026-07-29 · arXiv - AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching 🆕
Yiping Song, Jiaoyan Chen, Renate Schmidt et al. · 2026-07-29 · arXiv - On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment 🆕
Yongjian Guo, Wanlun Ma, Lingyu Shen et al. · 2026-07-29 · arXiv - Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data 🆕
Lingyang Zeng, Guangze Chen, Kaichen Yu et al. · 2026-07-29 · arXiv - Belief-Guided Decision Making with Uncertainty Gating in the Game of Go 🆕
Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand et al. · 2026-07-29 · arXiv - What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation 🆕
Vishisht Choudhary, Lukas Schmidt, Anne Zoë Kenntner et al. · 2026-07-29 · arXiv - From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence 🆕
Jia Luo · 2026-07-29 · arXiv - EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks 🆕
Peng Yin, Kai Li, Yifan Zhang et al. · 2026-07-29 · arXiv
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 🆕
Hengyi Xie, Chenfei Yao, Xianjin Wu et al. · 2026-07-29 · Hugging Face, arXiv · ▲109 - HumanCLAW: Can Vision-Language Models Act Through a Body? 🆕
Siyao Li, Jiawei Gu, Shuai Liu et al. · 2026-07-29 · Hugging Face, arXiv · ▲59 - StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation 🆕
Zijun Lin, Zeqing Wang, Cheston Tan et al. · 2026-07-29 · Hugging Face, arXiv · ▲10 - Anatomy Contextualized Adaption of CT Foundation Models 🆕
Roshan Kenia, Stephanie L McNamara, William Lotter · 2026-07-29 · arXiv - SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context 🆕
Zihan Deng, Chuanzhi Xu, Huiqi Liang et al. · 2026-07-29 · arXiv - Visual Credit Audit for Multimodal Spatial Reasoning 🆕
Feixiang Liu, Qiang Qiu, Lanbo Sun et al. · 2026-07-29 · arXiv - SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence 🆕
Chuanzhi Xu, Zihan Deng, Huiqi Liang et al. · 2026-07-29 · arXiv - ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection 🆕
Paul Julius Kühn, Saptarshi Neil Sinha, Tiago Kleist et al. · 2026-07-29 · arXiv - Progressive Multimodal Alignment for Continual Instruction Tuning 🆕
Duzhen Zhang, Yahan Yu, Qiaoyi Su et al. · 2026-07-29 · arXiv - Hearsay: Vision-Language Medical Diagnoses Without an Image 🆕
Siddharth Vohra · 2026-07-29 · arXiv - Mitigating Compounding Error via Video Representation Regularization 🆕
Taiye Chen, Qi Zhang, Yisen Wang · 2026-07-29 · arXiv - VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion 🆕
Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys · 2026-07-29 · arXiv - +25 more in this topic — see
data/papers.json
-
DLAM: Distributional Latent Actions with Temporal Constraints 🆕
Zuojin Tang, Feifan Luo, Haoyun Liu et al. · 2026-07-29 · arXiv -
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception 🆕
Gabe Everett, Brice Gunter, Ryan Vander Stelt et al. · 2026-07-29 · arXiv -
BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories 🆕
Zhe Liu, Quan Lu, Zhaohui Du et al. · 2026-07-29 · arXiv -
From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching 🆕
Shiyong Meng, Bolei Chen, Ping Zhong et al. · 2026-07-29 · arXiv -
Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights 🆕
Abhinav Sharma, Md Abdullah Al Hasan, Danjue Chen et al. · 2026-07-29 · arXiv -
RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models 🆕
Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer et al. · 2026-07-29 · arXiv -
Dense Soft Weighting for Radar Ego-Velocity Estimation 🆕
Atar Babgei, Chenyu Zhao, Michael Breza et al. · 2026-07-29 · arXiv -
NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and Education 🆕
Koneshka Bandyopadhyay, Ansh Mehta, Bassel El Mabsout et al. · 2026-07-29 · arXiv -
Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations 🆕
Jialiang Li, Yuhan Wang, Haojun Li et al. · 2026-07-29 · arXiv -
Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA 🆕
Tianhang Yang, Yanze Zheng, Junjie Wang et al. · 2026-07-29 · arXiv -
Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment 🆕
Marc Kaufeld, Dian Zhuang, Johannes Betz · 2026-07-29 · arXiv -
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation 🆕
Yushan Liu, Peibo Sun, Xintao Chao et al. · 2026-07-29 · arXiv - +13 more in this topic — see
data/papers.json
- The Sparsity Ceiling: Where Spiking Networks Can and Cannot Trade Activity for Energy 🆕
Zeyu Wang · 2026-07-29 · arXiv - Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression 🆕
Manuel Rodriguez · 2026-07-29 · arXiv - Reconstructing Backpropagation from Forward Fluctuations in Noise-modulated Neural Networks 🆕
Shuhei Ikemoto · 2026-07-29 · arXiv
- Cross-organisational Process Mining from Message Logs 🆕
Pieter Kwantes, Jetty Kleijn · 2026-07-29 · arXiv - DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution 🆕
Hanghui Guo, Weijie Shi, Zhangze Chen et al. · 2026-07-29 · arXiv - Living-Harness Is an Interactive-Agent Evolver 🆕
Yuetian Du, Yucheng Wang, He Xu et al. · 2026-07-29 · arXiv - One Run Is Not an Idea: The Implementation Lottery in Automated Research 🆕
Jingjie Ning, Shanshan Zhong, Xiaochuan Li et al. · 2026-07-29 · arXiv - A Persona-based Rate Action Index 🆕
Hayden Helm, Andrew Dassori · 2026-07-29 · arXiv - Towards a Systems Foundation for Agentic Cloud Management 🆕
Minghao Li, Ziqian Liu, Ziyu Mao et al. · 2026-07-28 · arXiv - CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents 🆕
Jiarun Fu, Lizhong Ding, Sida Chen et al. · 2026-07-28 · arXiv - ARCHER: Agentic Rule and Compliance Harness for Executable Regulations 🆕
Chiraag Singh Anand, Xue Wen Tan, Lionel Teo et al. · 2026-07-28 · arXiv - CoRenew: A large language model agent-based policy simulation platform for multifamily residential redevelopment 🆕
Yudi Zhang, Yuming Lin, Li Tian et al. · 2026-07-28 · arXiv - SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems 🆕
Haowen Dai, Zonghao Ying, Wenfeng Li et al. · 2026-07-28 · arXiv - Agentic AI-enabled discovery across large-scale sleep physiology 🆕
Rahul Thapa, Umaer Hanif, Robin Guillard et al. · 2026-07-28 · arXiv
- Improving Item Discoverability in e-Commerce Search via Related Intent Generation 🆕
Ji Xin, Xiao Xiao, Ishan Bhatt et al. · 2026-07-29 · arXiv - MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities 🆕
Francesco Tosoni · 2026-07-29 · arXiv - KAMR: Grounding Generation via Knowledge-Aligned Multi-hop Retrieval 🆕
Xiaochen Wang, Yuan Zhong, Haoyu Wang et al. · 2026-07-29 · arXiv - Learning from the Future: Privileged Self-Distillation for Sequential Recommendation 🆕
Jiakai Tang, Yang Zhang, See-Kiong Ng et al. · 2026-07-29 · arXiv - IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation 🆕
Yuheng Zheng, Yu Cui, Bin Wu et al. · 2026-07-29 · arXiv - Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising 🆕
Zipeng Chen, Jiaer Zheng, Xiangyang Xu et al. · 2026-07-29 · arXiv - CaIRec: Calibrated Modality Imputation for Incomplete Multimodal Recommendation 🆕
Ruiyu Liu, Xiaohao Liu, Miaomiao Cai et al. · 2026-07-29 · arXiv - WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models 🆕
Hao Jiang, Peiru Du, Pengfei Yao et al. · 2026-07-29 · arXiv - ASARL: Autonomous Social-Aware Relevance Learning for QQ Search 🆕
Tao Su, Jinjing Hu, Xiao Wang et al. · 2026-07-29 · arXiv - Multi-Decoder OneRec: Controllable Generative Retrieval for Multi-Objective Industrial Recommendation 🆕
You Wang, Zhao Liu, Guoping Tang et al. · 2026-07-29 · arXiv - NMKFR: A Robust Framework for Time-Aware Cold-Start Recommendation 🆕
Chengzhi Liu, Ning Zeng, Zehui Qu · 2026-07-29 · arXiv - PSG: Pair-Space Generation for Efficient Generative Reranking 🆕
Chao Feng, Li Ma, Xiancheng Gao et al. · 2026-07-29 · arXiv - +13 more in this topic — see
data/papers.json
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents 🆕
Amirmohammad Farzaneh, Osvaldo Simeone · 2026-07-29 · arXiv - PIKS: Universal Physics-Informed Kernel Methods 🆕
Joachim Bona-Pellissier, Giacomo Meanti, Matteo Santacesaria et al. · 2026-07-29 · arXiv - Feature Bagging Provides Stability 🆕
Yuheng Ma, Qiang Sun · 2026-07-29 · arXiv - Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions 🆕
Shuo-Chieh Huang, Chien-Ming Chi, Jau-er Chen · 2026-07-29 · arXiv - Conformalized Rate-Adaptive Sensing 🆕
Jiawei Yang, Yao Zhang · 2026-07-29 · arXiv - Crossing-Free Probabilistic K-Line Forecasts Without Retraining 🆕
Runyao Yu, Yuchen Tao, Yujie Chen et al. · 2026-07-29 · arXiv - Early Failure Prediction from Near-Anomaly Detection: A Proactive Approach 🆕
Léa Billet, Louise Travé-Massuyès, Elodie Chanthery et al. · 2026-07-29 · arXiv - Chaos Is a LADDER: Domain Generalization Beyond Invariance via Reweighting 🆕
Yuhang Jiang, Fengchuan Zhang, Sanguo Zhang et al. · 2026-07-29 · arXiv - Origins and mitigation of double descent in reduced order modeling 🆕
Andrei A. Klishin, J. Nathan Kutz, Krithika Manohar · 2026-07-29 · arXiv
- MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning 🆕
Weijie Wu, Junbo Li, Lin Li et al. · 2026-07-29 · arXiv
- The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making 🆕
Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos et al. · 2026-07-29 · arXiv - Human diversity fuels collective creativity that large language models cannot simulate or sustain 🆕
Mengchen Dong, Hiromu Yakura · 2026-07-29 · arXiv - AI as Friction for Reflection Support in Ideation 🆕
Janin Koch, Xiaohan Liao, Géry Casiez · 2026-07-29 · arXiv - FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets 🆕
Peisen Xu, Jérémie Garcia, Peter Cleveland et al. · 2026-07-29 · arXiv - Sensor-Placement-Agnostic Sonomyography: Toward Continuous High-Dimensional Control by Users with Tetraplegia 🆕
Gavin Sueltz, Vikram Athithan, Emma Ferran et al. · 2026-07-29 · arXiv
- How enabling two settings tripled our scores on the ARC-AGI-3 benchmark 🆕
· 2026-07-29 · OpenAI - Accelerating scientific discovery with ChatGPT for Academic Researchers 🆕
· 2026-07-29 · OpenAI - How GPT-5.6 fuses frontier intelligence with frontier efficiency 🆕
· 2026-07-29 · OpenAI
- SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution 🆕
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu et al. · 2026-07-29 · Hugging Face · ▲15 - MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair 🆕
Xuanze Chen, Xukang Xie, Wentao Fu et al. · 2026-07-29 · arXiv - Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning 🆕
Hongliang Zhang, Zhongyuan Yu, Guijuan Wang et al. · 2026-07-29 · arXiv - A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities 🆕
Wenhao Yang, Runzhi He, Minghui Zhou · 2026-07-29 · arXiv - MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis 🆕
Yihao Chen, Shi Chang, Khaled Chawa et al. · 2026-07-29 · arXiv - Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devices 🆕
Harvey Samuel George Johnson, Sendy Phang · 2026-07-29 · arXiv - Detecting seizure onset and offset times using human intelligence: A critical-transitions-based approach 🆕
Andrew Flynn, Cian McCafferty, Klaus Lehnertz et al. · 2026-07-29 · arXiv - InferScale: GPU-Native KV Injection for Personalized LLM Serving 🆕
Peter Li, Prashant Pandey · 2026-07-29 · arXiv - Field Codes for Distributed Coupling Samplers and Certified Empirical Transport 🆕
Hung Mai, Hai Nguyen, Luong Doan et al. · 2026-07-29 · arXiv - GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding 🆕
Jiale Chen, Torsten Hoefler, Dan Alistarh · 2026-07-29 · arXiv - HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models 🆕
Petr Simecek, Elnaz Babayeva, Jiri Balhar et al. · 2026-07-29 · arXiv - On the robustness of noisy solutions in non-convex neural networks 🆕
Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina · 2026-07-29 · arXiv - +41 more in this topic — see
data/papers.json
A scheduled GitHub Action runs scripts/update_papers.py, which pulls the newest submissions from arXiv (across the categories in config.yaml) and the community-curated Hugging Face Daily Papers, merges them, keeps a rolling 30-day window, and regenerates this file.
Edit config.yaml to change the arXiv categories, the window size, or the sources table — no code changes needed. See SETUP.md for how the automation is enabled.