Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

10 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

📄 AI Research Papers — Live Tracker

An auto-updating feed of the newest AI research, gathered from the top sources and refreshed on a schedule. Freshly added papers are marked 🆕.

Papers Window Updated

250 papers from the last 30 days · Last updated 2026-07-30 12:09 UTC

🔗 Top Sources for AI Research

Where the papers below come from, plus the major labs and venues worth watching.

Source What it is
arXiv (cs.AI / cs.LG / cs.CL / cs.CV) The primary open preprint server — where nearly all AI research lands first. (Auto-tracked)
Hugging Face Papers Community-curated daily & trending papers. Official successor to Papers with Code. (Auto-tracked)
OpenReview Open peer review for ICLR, NeurIPS, and other top venues.
Semantic Scholar AI-powered scholarly search with citation graph and a free API.
OpenAlex Fully open index of the world's scholarly papers; free API.
Google DeepMind — Research Publications from one of the leading industrial AI labs. (Watch — no reliable public feed yet)
OpenAI — Research Research index and technical reports from OpenAI. (Auto-tracked via RSS)
Anthropic — Research Papers on safety, interpretability, and frontier models. (Auto-tracked via scraping — best-effort)
Meta AI (FAIR) — Research Fundamental AI Research publications from Meta. (Auto-tracked via scraping — best-effort)
Google Research Broad publications across ML, systems, and applications. (Auto-tracked via RSS)
Microsoft Research Extensive AI and systems research output.
Papers with Code (archived) Retired by Meta in July 2025; historical SOTA/benchmark data is frozen here.

⭐ Highlights

Why these are worth your time — picked from the top of the feed.



🔥 Trending (Hugging Face community picks)

▲ Upvotes Paper Date
140 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone 🆕 2026-07-28
109 TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 🆕 2026-07-29
70 CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents 🆕 2026-07-28
60 CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization 🆕 2026-07-28
59 HumanCLAW: Can Vision-Language Models Act Through a Body? 🆕 2026-07-29
58 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition 🆕 2026-07-28
51 DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space 🆕 2026-07-28
38 CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition 🆕 2026-07-28
28 Pass the Baton: Trajectory-Relayed On-Policy Distillation 🆕 2026-07-28
27 CAST: Game Solvers as Turn-Level Teachers for LLM Agents 🆕 2026-07-28
17 Wonder: Video World Model Done Better 🆕 2026-07-28
15 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution 🆕 2026-07-29
14 Visual prompt engineering for video models 🆕 2026-07-28
14 Shieldstral 🆕 2026-07-28
12 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities 🆕 2026-07-28

🗓️ Released in the Last 30 Days

The newest papers (80 shown), most recent first — each with a direct link and a short description.

2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲109
Hengyi Xie, Chenfei Yao, Xianjin Wu et al.

Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial…

🔗 https://arxiv.org/abs/2607.27205

2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲59
Siyao Li, Jiawei Gu, Shuai Liu et al.

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed…

🔗 https://arxiv.org/abs/2607.27180

2026-07-29 · ✨ Other AI Research · Hugging Face · ▲15
Zhiyuan Yao, Yuxin Chen, Zhengxi Lu et al.

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one ta…

🔗 https://huggingface.co/papers/2607.26784

2026-07-29 · 👁️ Computer Vision · Hugging Face, arXiv · ▲10
Zijun Lin, Zeqing Wang, Cheston Tan et al.

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activat…

🔗 https://arxiv.org/abs/2607.26754

2026-07-29 · 🤖 Artificial Intelligence (General) · Hugging Face, arXiv · ▲6
Jingbo Zhou, Yusai Zhao, Qi Bao et al.

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a bench…

🔗 https://arxiv.org/abs/2607.27155

2026-07-29 · 🤖 Artificial Intelligence (General) · Hugging Face, arXiv · ▲5
Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generat…

🔗 https://arxiv.org/abs/2607.27191

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Julien Benchek, Austin Bennett, Jasmin Kern et al.

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set com…

🔗 https://arxiv.org/abs/2607.27189

2026-07-29 · 🧑‍💻 Human-AI Interaction · arXiv
Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos et al.

Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analy…

🔗 https://arxiv.org/abs/2607.27179

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Peter Tisnikar, Maja Swieczkowska, Benteng Ma et al.

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute…

🔗 https://arxiv.org/abs/2607.27177

2026-07-29 · 🔎 Information Retrieval & RAG · arXiv
Ji Xin, Xiao Xiao, Ishan Bhatt et al.

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the di…

🔗 https://arxiv.org/abs/2607.27172

2026-07-29 · 👁️ Computer Vision · arXiv
Roshan Kenia, Stephanie L McNamara, William Lotter

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning ana…

🔗 https://arxiv.org/abs/2607.27154

2026-07-29 · 🧠 Machine Learning · arXiv
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage gua…

🔗 https://arxiv.org/abs/2607.27143

2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Zuojin Tang, Feifan Luo, Haoyun Liu et al.

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations w…

🔗 https://arxiv.org/abs/2607.27138

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Suhas Thejaswi, Juhi Kulshreshta, Lutz Oettershagen

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduc…

🔗 https://arxiv.org/abs/2607.27134

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Yiping Song, Jiaoyan Chen, Renate Schmidt et al.

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we…

🔗 https://arxiv.org/abs/2607.27130

2026-07-29 · 🔊 Audio & Speech · arXiv
Weijie Wu, Junbo Li, Lin Li et al.

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to d…

🔗 https://arxiv.org/abs/2607.27109

2026-07-29 · 👁️ Computer Vision · arXiv
Zihan Deng, Chuanzhi Xu, Huiqi Liang et al.

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs…

🔗 https://arxiv.org/abs/2607.27084

2026-07-29 · 🧠 Machine Learning · arXiv
Yicheng Feng, Yan Zhang, Yan Cheng et al.

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and priv…

🔗 https://arxiv.org/abs/2607.27083

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Yongjian Guo, Wanlun Ma, Lingyu Shen et al.

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values…

🔗 https://arxiv.org/abs/2607.27081

2026-07-29 · ✨ Other AI Research · arXiv
Xuanze Chen, Xukang Xie, Wentao Fu et al.

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent ben…

🔗 https://arxiv.org/abs/2607.27080

2026-07-29 · 🧠 Machine Learning · arXiv
Vaneet Aggarwal

We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret…

🔗 https://arxiv.org/abs/2607.27073

2026-07-29 · 👁️ Computer Vision · arXiv
Feixiang Liu, Qiang Qiu, Lanbo Sun et al.

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decisio…

🔗 https://arxiv.org/abs/2607.27069

2026-07-29 · 👁️ Computer Vision · arXiv
Chuanzhi Xu, Zihan Deng, Huiqi Liang et al.

Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image…

🔗 https://arxiv.org/abs/2607.27066

2026-07-29 · 👁️ Computer Vision · arXiv
Paul Julius Kühn, Saptarshi Neil Sinha, Tiago Kleist et al.

While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale anno…

🔗 https://arxiv.org/abs/2607.27065

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Lingyang Zeng, Guangze Chen, Kaichen Yu et al.

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However,…

🔗 https://arxiv.org/abs/2607.27056

2026-07-29 · 🧠 Machine Learning · arXiv
Fengming Yu, Haiwei Pan, Kejia Zhang et al.

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences…

🔗 https://arxiv.org/abs/2607.27054

2026-07-29 · 🧠 Machine Learning · arXiv
Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet et al.

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requi…

🔗 https://arxiv.org/abs/2607.27023

2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Gabe Everett, Brice Gunter, Ryan Vander Stelt et al.

Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations t…

🔗 https://arxiv.org/abs/2607.26985

2026-07-29 · 👁️ Computer Vision · arXiv
Duzhen Zhang, Yahan Yu, Qiaoyi Su et al.

Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it central to cross-modal understanding. In Multimodal Continual Instruction Tuning (MCIT), however, shifting visual distributions and evolving…

🔗 https://arxiv.org/abs/2607.26947

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand et al.

Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-gra…

🔗 https://arxiv.org/abs/2607.26946

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Vishisht Choudhary, Lukas Schmidt, Anne Zoë Kenntner et al.

Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection f…

🔗 https://arxiv.org/abs/2607.26935

2026-07-29 · ✨ Other AI Research · arXiv
Hongliang Zhang, Zhongyuan Yu, Guijuan Wang et al.

Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthi…

🔗 https://arxiv.org/abs/2607.26933

2026-07-29 · 🦾 Robotics & Embodied AI · arXiv
Zhe Liu, Quan Lu, Zhaohui Du et al.

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms are designed for household environments and treat a target as an object center or an arbitrary nearby position. This representation is inad…

🔗 https://arxiv.org/abs/2607.26914

2026-07-29 · 🧠 Machine Learning · arXiv
Brandon Gower-Winter, Georg Krempl

In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they predict. This phenomena is known as Outcome Performativity. This paper formalises an approach for detecting Outcome Performativity using predi…

🔗 https://arxiv.org/abs/2607.26908

2026-07-29 · 🤖 Artificial Intelligence (General) · arXiv
Jia Luo

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource f…

🔗 https://arxiv.org/abs/2607.26903

2026-07-29 · 🧑‍💻 Human-AI Interaction · arXiv
Mengchen Dong, Hiromu Yakura

Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges…

🔗 https://arxiv.org/abs/2607.26899

2026-07-29 · 👁️ Computer Vision · arXiv
Siddharth Vohra

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and derm…

🔗 https://arxiv.org/abs/2607.26886

2026-07-29 · 📈 ML Theory & Statistics · arXiv
Amirmohammad Farzaneh, Osvaldo Simeone

LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, and control of physical AI systems. Yet, when deployed at the edge, they must tightly manage their reasoning budget while remainin…

🔗 https://arxiv.org/abs/2607.26865

2026-07-29 · 🧠 Machine Learning · arXiv
Junoh Park, Junseo Hwang, Wonguk Cho et al.

Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage…

🔗 https://arxiv.org/abs/2607.26862

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Ruikang Zhang, Shuo Wang, Qi Su

Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have prima…

🔗 https://arxiv.org/abs/2607.26853

2026-07-29 · 🧠 Machine Learning · arXiv
Yansen Zhang, Yilu Liu, Tianyu Liu et al.

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Existing adaptive discovery controllers assign credit based only on score progress, even though prompt length, retries, and guidance calls cau…

🔗 https://arxiv.org/abs/2607.26828

2026-07-29 · 🧑‍💻 Human-AI Interaction · arXiv
Janin Koch, Xiaohan Liao, Géry Casiez

Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster output translate into more value for the designer. We argue, however, that this framing leaves out something important about how d…

🔗 https://arxiv.org/abs/2607.26827

2026-07-29 · ✨ Other AI Research · arXiv
Wenhao Yang, Runzhi He, Minghui Zhou

Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate coding agents' behavior, spanning from a total ban, mandatory disclosure, to verification gates and human sign-offs. Yet, whether coding agents…

🔗 https://arxiv.org/abs/2607.26819

2026-07-29 · 🧠 Machine Learning · arXiv
Zhaoyang Ma, Zhihao Wu, Xin Gao et al.

Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local architectures often induce non-aligned representation spaces, making it difficult to transfer global knowledge across silos. Existing…

🔗 https://arxiv.org/abs/2607.26801

2026-07-29 · 🧠 Machine Learning · arXiv
Perry Dong, Ron Polonsky, Dorsa Sadigh et al.

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional w…

🔗 https://arxiv.org/abs/2607.27203

2026-07-29 · 🧠 Machine Learning · arXiv
Shady E. Ahmed, Panos Stinis

We present a novel approach to regression tasks using classification which is motivated by the mechanism used by fruitflies to sense their environment. Specifically, we formulate a general framework for learning nonlinear input-output relationships by replacing complex global su…

🔗 https://arxiv.org/abs/2607.27196

2026-07-29 · 🧠 Machine Learning · arXiv
Lennon J. Shikhman, Michael Galarnyk, Aadi Dash et al.

Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only…

🔗 https://arxiv.org/abs/2607.27188

2026-07-29 · 🧠 Machine Learning · arXiv
Quang Bui, Sparsh Roy, Akash Gundimeda et al.

Solving a continuous algebraic constraint system requires two decisions: which values satisfy the constraints, and which structural augmentation renders an unsolvable system solvable. Classical solvers answer the first well and the second only by enumeration. On that discrete de…

🔗 https://arxiv.org/abs/2607.27169

2026-07-29 · 🧠 Machine Learning · arXiv
Francesco Pinto, Luca Lanzilao, Paco Lopez Dekker et al.

Accurate intraday forecasts of offshore wind are becoming increasingly important for power system operation and the integration of growing shares of offshore wind energy. Operational forecasts rely predominantly on numerical weather prediction (NWP), which is not optimized for l…

🔗 https://arxiv.org/abs/2607.27152

2026-07-29 · ✨ Other AI Research · arXiv
Yihao Chen, Shi Chang, Khaled Chawa et al.

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on Prog…

🔗 https://arxiv.org/abs/2607.27146

2026-07-29 · ✨ Other AI Research · arXiv
Harvey Samuel George Johnson, Sendy Phang

This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic. With the target workload of branch prediction (BP) for multistage pipelined central pro-cessing unit (CPU) cores. For thi…

🔗 https://arxiv.org/abs/2607.27140

2026-07-29 · 🧠 Machine Learning · arXiv
Zuyuan Zhang, Yongshan Chen, Mahdi Imani et al.

An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision pro…

🔗 https://arxiv.org/abs/2607.27132

2026-07-29 · 🧠 Machine Learning · arXiv
Kaifeng Zhang, Kai Ming Ting

Persistence Diagram (PD) is known to capture point cloud topology effectively, but its computation has high time complexity. Expected Persistence Diagram (EPD) has been developed to reduce the time cost by studying the topology of multiple subsets of a point cloud and it serves…

🔗 https://arxiv.org/abs/2607.27126

2026-07-29 · 🧠 Machine Learning · arXiv
Filipa Lino, Bárbara Tavares, Carlos Santiago et al.

Emergency Departments (EDs) are critical access points in healthcare systems, yet they face persistent pressure from unpredictable patient demand, seasonal surges, and non-urgent visits. Effective ED planning requires forecasts at multiple decision-making levels: hospitals need…

🔗 https://arxiv.org/abs/2607.27106

2026-07-29 · ✨ Other AI Research · arXiv
Andrew Flynn, Cian McCafferty, Klaus Lehnertz et al.

Most existing seizure detection algorithms require extensive pre-processing of the data and rely on heuristic or currently unexplainable machine learning approaches. These approaches often struggle with balancing detection sensitivity and specificity in the presence of variable…

🔗 https://arxiv.org/abs/2607.27105

2026-07-29 · 🧠 Machine Learning · arXiv
Aleksandr Berdnikov, Yevgeny Liokumovich

We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. We find that most of the considered open-source models do have such a representation, and it often even surfaces to the top principal compone…

🔗 https://arxiv.org/abs/2607.27092

2026-07-29 · ✨ Other AI Research · arXiv
Peter Li, Prashant Pandey

Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests. Production memory systems (e.g., Mem0, MemGPT, and Zep) retrieve a relevant sub…

🔗 https://arxiv.org/abs/2607.27090

2026-07-29 · ✨ Other AI Research · arXiv
Hung Mai, Hai Nguyen, Luong Doan et al.

In this paper, we formulate three communication tasks for empirical optimal transport: distributed coupling sampling, cost-evaluable coupling output, and scalar value-certified sampling. Our main result is a field-code compiler: any communicated transport field approximating an…

🔗 https://arxiv.org/abs/2607.27078

2026-07-29 · 🧠 Machine Learning · arXiv
Nicolas Béreux, Aurélien Decelle, Cyril Furtlehner et al.

Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the c…

🔗 https://arxiv.org/abs/2607.27077

2026-07-29 · 🧠 Machine Learning · arXiv
Kindeep K. Dhatt, Tengyue Wu, Hanbang Hua et al.

Continuous cuffless blood pressure (BP) monitoring remains challenging due to motion artifacts, physiological variability, and the limited robustness of conventional pulse transit time (PTT) models under dynamic conditions. Many prior approaches rely on multi-second windows to s…

🔗 https://arxiv.org/abs/2607.27076

2026-07-29 · 📈 ML Theory & Statistics · arXiv
Joachim Bona-Pellissier, Giacomo Meanti, Matteo Santacesaria et al.

Physics-informed machine learning incorporates physical principles --often expressed via differential operators-- into data-driven models. While physics-informed neural networks (PINNs) dominate empirical applications, the complexity of neural network architectures and optimizat…

🔗 https://arxiv.org/abs/2607.27062

2026-07-29 · ✨ Other AI Research · arXiv
Jiale Chen, Torsten Hoefler, Dan Alistarh

Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. They process the entries in a fixed order, one at a time, propagating each rounding error to the entries not yet processed through a…

🔗 https://arxiv.org/abs/2607.27042

2026-07-29 · 👁️ Computer Vision · arXiv
Taiye Chen, Qi Zhang, Yisen Wang

Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomeno…

🔗 https://arxiv.org/abs/2607.27036

2026-07-29 · 🧠 Machine Learning · arXiv
Bum Jun Kim

Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. Moreover, prior work has not considered actual deployment conditions, wh…

🔗 https://arxiv.org/abs/2607.27031

2026-07-29 · ✨ Other AI Research · arXiv
Petr Simecek, Elnaz Babayeva, Jiri Balhar et al.

LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark…

🔗 https://arxiv.org/abs/2607.27030

2026-07-29 · 🧠 Machine Learning · arXiv
James Chapman

Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-a…

🔗 https://arxiv.org/abs/2607.27027

2026-07-29 · 🧠 Machine Learning · arXiv
Kaizhen Tan, Xin Xu, Siru Tao et al.

A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEW…

🔗 https://arxiv.org/abs/2607.27017

2026-07-29 · ✨ Other AI Research · arXiv
Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina

Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare.…

🔗 https://arxiv.org/abs/2607.27000

2026-07-29 · ✨ Other AI Research · arXiv
Ruoyu Wang, Heng Zhao, Renjie Wu et al.

Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. H…

🔗 https://arxiv.org/abs/2607.26998

2026-07-29 · 🧠 Machine Learning · arXiv
Peter Lorenz, Anjith George, Sébastien Marcel

Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained vision models rather than training task-specific ar…

🔗 https://arxiv.org/abs/2607.26993

2026-07-29 · ✨ Other AI Research · arXiv
Franz Nowak, Ryan Cotterell, Reda Boumasmoud

What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length? Existing answers often rely on idealized arithmetic, but under finite precision, rounding and evaluation order can change what information attention retains…

🔗 https://arxiv.org/abs/2607.26988

2026-07-29 · ✨ Other AI Research · arXiv
Michelangelo Domina, Michele Ceriotti

Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc.,…

🔗 https://arxiv.org/abs/2607.26984

2026-07-29 · 📈 ML Theory & Statistics · arXiv
Yuheng Ma, Qiang Sun

We study feature bagging through the lens of algorithmic stability. Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner. We introduce feature instability (FI), the feature-axis…

🔗 https://arxiv.org/abs/2607.26964

2026-07-29 · 📈 ML Theory & Statistics · arXiv
Shuo-Chieh Huang, Chien-Ming Chi, Jau-er Chen

Minimax-optimal rates for multivariate distribution estimation are known to suffer from the curse of dimensionality. We propose a sparse Bayesian network approach in which each conditional probability is estimated using sparsity-aware conditional mean methods. The resulting esti…

🔗 https://arxiv.org/abs/2607.26955

2026-07-29 · 🧠 Machine Learning · arXiv
Alexandr Udeneev, Petr Babkin, Oleg Bakhteev

Ensembles are a standard way to improve the performance and robustness of deep neural networks, but their effectiveness crucially depends on both the quality and the diversity of individual models. Most neural architecture search (NAS) methods are computationally expensive. Exte…

🔗 https://arxiv.org/abs/2607.26940

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Hao Fei, Yiran Zhao

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, an…

🔗 https://arxiv.org/abs/2607.27201

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Ben Glickenhaus, Katherine Thai, Jenna Russell et al.

We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, P…

🔗 https://arxiv.org/abs/2607.27183

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Raphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior et al.

State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first…

🔗 https://arxiv.org/abs/2607.27178

2026-07-29 · ✨ Other AI Research · arXiv
Yihao Chen, Shi Chang, Feng Lin et al.

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an exec…

🔗 https://arxiv.org/abs/2607.27167

2026-07-29 · 🗣️ NLP & Large Language Models · arXiv
Jiayuan Di, Haoyi Yang, Yufei Luo et al.

Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce…

🔗 https://arxiv.org/abs/2607.27022


📑 Topics

  • 🗣️ NLP & Large Language Models (29)
  • 🧠 Machine Learning (38)
  • 🤖 Artificial Intelligence (General) (11)
  • 👁️ Computer Vision (37)
  • 🦾 Robotics & Embodied AI (25)
  • 🧬 Neural & Evolutionary Computing (3)
  • 🤝 Multi-Agent Systems (11)
  • 🔎 Information Retrieval & RAG (25)
  • 📈 ML Theory & Statistics (9)
  • 🔊 Audio & Speech (1)
  • 🧑‍💻 Human-AI Interaction (5)
  • 🏛️ From the Labs (blogs & reports) (3)
  • ✨ Other AI Research (53)

🗣️ NLP & Large Language Models (29)

🧠 Machine Learning (38)

🤖 Artificial Intelligence (General) (11)

👁️ Computer Vision (37)

🦾 Robotics & Embodied AI (25)

🧬 Neural & Evolutionary Computing (3)

🤝 Multi-Agent Systems (11)

🔎 Information Retrieval & RAG (25)

📈 ML Theory & Statistics (9)

🔊 Audio & Speech (1)

🧑‍💻 Human-AI Interaction (5)

🏛️ From the Labs (blogs & reports) (3)

✨ Other AI Research (53)


🤖 How this works

A scheduled GitHub Action runs scripts/update_papers.py, which pulls the newest submissions from arXiv (across the categories in config.yaml) and the community-curated Hugging Face Daily Papers, merges them, keeps a rolling 30-day window, and regenerates this file.

Edit config.yaml to change the arXiv categories, the window size, or the sources table — no code changes needed. See SETUP.md for how the automation is enabled.

About

A self-updating feed of the latest AI research papers from arXiv and Hugging Face, organized by topic.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages