Safety in Self-Evolving Agents: A Survey
arXiv:2610.00093v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from...
View ArticleCharacterizing and Codifying Malware Sophistication
arXiv:2610.00098v1 Announce Type: new Abstract: 'Sophisticated' is widely used to describe malware, yet it lacks a consistent definition within academic literature. While existing software quality and...
View ArticleA Comprehensive Review of One-Pixel Attack: Research Status, Taxonomy,...
arXiv:2610.00125v1 Announce Type: new Abstract: One-Pixel Attacks (OPAs) represent one of the most extreme demonstrations of adversarial fragility in deep learning, where modifying a single pixel can...
View ArticleA Verifier Can Leak the Answer: Diagnosability Before Optimization in...
arXiv:2610.00126v1 Announce Type: new Abstract: Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can...
View ArticleThe Cognitive Continuity Test: Verifying Governed State Transitions in...
arXiv:2610.00132v1 Announce Type: new Abstract: Persistent AI agents revise beliefs, consolidate memory, and replace execution substrates. Similar successor states can accompany differently authorized...
View ArticleEvasion Attacks: How Adversarial Noise Bypasses ML Classifiers
arXiv:2610.00136v1 Announce Type: new Abstract: This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network...
View ArticleIntrusion Detection for Agentic Processes: Evidence-Based Runtime Monitoring
arXiv:2610.00151v1 Announce Type: new Abstract: Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval....
View ArticleTokenized Key-Gated Adapter Routing: A Secure Access Control Mechanism...
arXiv:2610.00309v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in privacy-critical domains (e.g., healthcare, finance, and government), but their propensity to...
View ArticleActions with Receipts: Jointly Binding Claims, Evidence, and Execution for...
arXiv:2610.00327v1 Announce Type: new Abstract: Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim...
View ArticleUnifiedAttack: Evaluating the Safety of Large Multimodal Models in...
arXiv:2610.00341v1 Announce Type: new Abstract: As Large Multimodal Models (LMMs) transition toward natively unified architectures, evaluating their safety in synergistic harmful image-text generation...
View ArticleAuthorization for Self-Modifying AI Agent Populations: Conserving Authority...
arXiv:2610.00347v1 Announce Type: new Abstract: Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does...
View ArticleRemoving the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight...
arXiv:2610.00348v1 Announce Type: new Abstract: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input....
View ArticleProof-Gated Signing: Solver-Checked Transaction Guards that Hold Under State...
arXiv:2610.00354v1 Announce Type: new Abstract: AI agents that control wallets read attacker-reachable content, so they can be steered into proposing harmful transactions. The usual last line of...
View ArticleOn the Relationship between Model Quantization and Model Inversion Attacks
arXiv:2610.00382v1 Announce Type: new Abstract: Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion...
View ArticleFrom A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM...
arXiv:2610.00392v1 Announce Type: new Abstract: Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task...
View ArticleZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning...
arXiv:2610.00450v1 Announce Type: new Abstract: Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as...
View ArticleHarbormaster: Evidence-Gated, Replay-Safe Maritime Anomaly Detection on AWS
arXiv:2610.00519v1 Announce Type: new Abstract: Ships broadcast their positions through the Automatic Identification System (AIS), and those reports can be false or missing. An operator who acts on an...
View ArticleNo One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical...
arXiv:2610.00557v1 Announce Type: new Abstract: Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement...
View ArticleTowards Hierarchical Cyber Defense with Large Language Models: From Planning...
arXiv:2610.00590v1 Announce Type: new Abstract: An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to...
View ArticleProgressive-Resolution Secure Aggregation for Federated Learning
arXiv:2610.00695v1 Announce Type: new Abstract: Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the...
View Article