NousResearch/hermes-agent 是一个开源的、可扩展的智能体框架,旨在随用户需求演进,支持动态工具调用与上下文自适应。该项目提供代码仓库、配置示例和基础文档,便于开发者快速集成与定制。 NousResearch/hermes-agent is an open-source, extensible agent framework designed to evolve with user needs, supporting dynamic tool use and context-aware adaptation. The repository includes code, configuration examples, and foundational documentation for developer integration and customization.
这是一个面向AI编程助手(如Claude Code、Codex等)的代理性能优化系统,聚焦于技能编排、本能建模、记忆机制、安全增强和以研究为先的开发范式。 This is an agent-centric performance optimization system designed for AI coding assistants (e.g., Claude Code, Codex, Opencode, Cursor), emphasizing skill orchestration, instinct modeling, memory integration, security, and research-first development.
Hugging Face Transformers 是一个广泛使用的开源库,提供数千种预训练模型和统一API,支持文本、视觉、音频及多模态任务的推理与训练。 Hugging Face Transformers is a widely adopted open-source library offering thousands of pre-trained models and a unified API for inference and training across text, vision, audio, and multimodal tasks.
本文探讨了AI在测试用例生成中的实际效能与局限性,指出其虽能快速生成看似合理的测试代码,但常因缺乏深层语义理解而产生无效或误导性断言。 This article critically examines AI's role in test generation, highlighting its speed and surface-level realism while warning about its tendency to produce syntactically plausible but semantically incorrect or ineffective tests.
该GitHub仓库是Ollama项目,提供本地运行多种主流开源大模型(如Qwen、Gemma、DeepSeek等)的轻量级工具,支持一键拉取与部署。 This GitHub repository is the Ollama project, a lightweight tool for running popular open-source LLMs (e.g., Qwen, Gemma, DeepSeek) locally with one-command setup and model management.
Dify 是一个面向生产环境的开源平台,用于构建和部署基于智能体(agentic)的工作流应用,支持可视化编排、模型集成与 API 发布。 Dify is a production-ready open-source platform for building and deploying agentic workflow applications, featuring visual orchestration, LLM integration, and API publishing.
vLLM 是一个高性能、内存高效的大型语言模型推理与服务引擎,专为提升吞吐量和降低显存占用而设计,广泛用于生产环境中的 LLM 部署。 vLLM is a high-throughput, memory-efficient inference and serving engine for large language models, designed to optimize throughput and reduce GPU memory usage in production deployments.
Immich 3.0 是一款开源自托管照片和视频管理工具的新版本,增强了AI驱动的搜索、人脸识别和跨设备同步功能。 Immich 3.0 is a new version of the open-source, self-hosted photo and video management tool, featuring enhanced AI-powered search, face recognition, and cross-device synchronization.
Kimi K2.7 Code模型已正式集成至GitHub Copilot,用户可在支持环境中直接调用该AI编程辅助功能。 The Kimi K2.7 Code model is now generally available within GitHub Copilot, enabling users to access this AI-powered coding assistant directly in supported environments.
Langflow 是一个开源的低代码平台,用于可视化构建、调试和部署基于大语言模型的AI代理与工作流。 Langflow is an open-source, low-code platform for visually designing, debugging, and deploying LLM-powered AI agents and workflows.
MoneyPrinterTurbo 是一个开源项目,利用大语言模型和多模态AI技术,实现一键生成高清短视频,显著简化视频创作流程。 MoneyPrinterTurbo is an open-source project that leverages large language models and multimodal AI to generate high-definition short videos with a single click, streamlining video creation workflows.
cc-switch 是一款跨平台桌面AI助手,集成Claude Code、Codex、OpenCode、OpenClaw、Gemini CLI和Hermes Agent等多种AI编程工具,提供统一界面访问。其官方网址为ccswitch.io。 cc-switch is a cross-platform desktop AI assistant that unifies access to multiple AI coding tools—including Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI, and Hermes Agent—via a single interface. The official website is ccswitch.io.
SkillCoach 是一种新型自我演化的评估框架,旨在精细化评估和提升大语言模型智能体对技能(如SOP、工具工作流)的使用能力,解决现有技能库中技能重叠与粗粒度验证导致的不可靠问题。 SkillCoach is a novel self-evolving rubric framework designed to finely evaluate and enhance how LLM agents select, compose, and execute reusable skills—addressing reliability issues caused by skill overlap and overly coarse final verification in real-world skill repositories.
本文提出PACE(Agentic Capability Evaluation Proxy),一种通过轻量级非代理式基准(如推理、代码生成)预测大型语言模型代理在昂贵代理基准(如SWE-Bench、GAIA)上表现的新评估范式,旨在显著降低代理能力评估的成本与复杂度。 This paper introduces PACE (Proxy for Agentic Capability Evaluation), a novel evaluation paradigm that predicts LLM agent performance on expensive, infrastructure-heavy agentic benchmarks (e.g., SWE-Bench, GAIA) using fast, low-cost non-agentic capability benchmarks (e.g., reasoning, code generation), enabling scalable and affordable agent evaluation.
WorldDirector 是一种新型视频世界模型框架,通过将语义运动编排与视觉生成解耦,并利用大语言模型协调3D轨迹与相机运动,实现持久动态对象记忆和自由视角探索。 WorldDirector is a novel video world model framework that explicitly decouples semantic motion orchestration from visual generation, enabling persistent dynamic object memory and unrestricted viewpoint exploration via LLM-coordinated 3D trajectories and camera control.
本文是一篇arXiv预印本论文,系统评估了新型矩阵结构优化器(SOAP、Muon及其混合体)在机器学习原子间势(MLIPs)训练中的性能,挑战了当前普遍采用Adam优化器的惯例。 This arXiv preprint systematically evaluates novel matrix-structured optimizers—SOAP, Muon, and their hybrid SOAP-Muon—for training machine learning interatomic potentials (MLIPs), challenging the field’s default reliance on Adam.
Manufact(YC S25)是一家新推出的云服务平台,专注于支持MCP(Model Control Protocol)应用与服务器部署,同时维护开源MCP SDK项目mcp-use。 Manufact (YC S25) is a newly launched cloud platform for MCP (Model Control Protocol) applications and servers, while continuing open-source SDK development under the mcp-use name.
LobeHub 是一个开源的 AI 代理编排平台,旨在将多个 AI 智能体组织为 7×24 小时持续运行的“AI 团队”,提供招聘、调度与汇报等自动化管理能力。 LobeHub is an open-source AI agent orchestration platform designed to organize multiple AI agents into a 24/7 operational 'AI team', offering automated capabilities for hiring, scheduling, and reporting.
本文提出DiscoBench——首个专注于澄清意识(clarification-aware)的深度搜索基准,旨在评估搜索代理在面对模糊、不完整或错误用户查询时主动提问与澄清的能力。 This paper introduces DiscoBench, the first benchmark specifically designed to evaluate search agents’ ability to proactively ask clarifying questions in deep search scenarios where user queries are vague, underspecified, or factually incorrect.
本文探讨了混合注意力模型中全注意力层的选择策略,指出现有启发式方法忽视层间依赖性,提出需建模层间协同关系以提升长上下文建模效率。 This work investigates layer selection strategies for hybrid attention models, arguing that existing heuristic methods (e.g., fixed patterns or isolated layer scoring) overlook inter-layer dependencies—critical for optimizing long-context efficiency in Transformer-to-hybrid conversion.