⚡ GitHub Pulse · 早报

生成于 2026-09-20 07:22 UTC · 追踪 2,560 仓库 · 8 多源共振
今日榜单几乎被“给 agent 用的工具”占据——浏览器操作、电脑操作、技能包、开发沙箱、自主研究——重心正从“做模型”转向“做 agent 基础设施层”。jev-ultrafast 等经核实为真实发布,新仓库 PR 少属正常,并非刷榜。
🔍 搜索中 · 显示所有标签页的匹配项 · 按 Esc 清除
📖 一句话:今天几乎所有热度都在“给 agent 用的工具”上——浏览器操作、电脑操作、技能包、沙箱,重心已从‘做模型’转到‘做 agent 的手脚和环境’。先看下面两条(jev-ultrafast 与 trycua/cua),它们代表 computer-use 的两条路线;模型榜变化不大、快讯里刷屏的 Muse 系列可快速扫过。注意:榜首的 jev 集群已核实是真实开源发布,不是刷榜——新仓库 PR 少属正常。

📌 必读 导读 · 今天先看这些

  • Jev Ultrafast: a browser agent with a dynamic, indexed action space (HN 讨论) HN
    浏览器 agent 的新范式:读 DOM 表格、免视觉模型、70–500ms 决策;评论区有一线工程判断,比 star 数更有信息量。
  • browser-use/jev-ultrafast repo
    直接读实现:indexed-DOM + speculative fan-out 如何把浏览器自动化压到 ~7s;MVP 局限(Shadow DOM/iframe/无结果校验)也写得很坦白。
  • trycua/cua — Computer-use 2.0 repo
    另一条路线:跨 OS 驱动 + agent 集群 + 评测/训练数据。想理解“agent 基础设施”全景,这个和 jev 对照着看。
  • hypit-ai/hypit repo
    把短视频还原成可编辑“代码”、批量产变体;也是“skill 即分发”(npx skills add)的样例——一个正在成形的能力分发范式。

🔬 深度洞察 deep research

Agent 基础设施成为重心:大家在造“手和沙箱”,不是又一个模型
今日多源共振里,7/9 是“让 agent 能操作/被隔离/被评测”的工具,而非模型本身。这说明生态的建设热度已从“更强的大脑”转向“可用的手脚与环境”——谁能让 agent 稳定操作真实系统,谁就卡住下一段价值。对开发者:该层仍很早、可切入;对投资人:这是当前最密集的在建方向。
browser-use/jev-ultrafasttrycua/cuaarcboxlabs/arcboxcoder/codercloudflare/security-audit-skillsapientinc/PRAXISTdeeplethe/utopia
“Skill(可安装技能包)”正成为能力分发的新范式
cloudflare/security-audit-skill 与 hypit 都以“技能包”形式分发(hypit 用 `npx skills add`),即把一项能力打包给 Claude Code/Codex 等 agent 直接调用。这是一个正在成形的分发 primitive——值得盯它会不会长出“agent 能力的应用商店”。
cloudflare/security-audit-skillhypit-ai/hypit
专用快模型分工化:浏览器 agent 成本正在坍塌
browser-use/jev-ultrafast 建立在 TypeSafe 的 Jev(“System One”)模型上——返回类型化概率决策而非文本、跳过视觉模型、70–500ms 出结果;自测把 Google Flights 自动化压到 7s、成本降约 90%(注:官方自测、未经第三方复现)。呼应 ai_news 里 GLM-5.3 FlashX 等“快变体”:趋势是工作流不同环节用不同的更快更便宜的模型,而非一个大模型通吃。
browser-use/jev-ultrafasttamaratran/fast-jev-compactionGLM-5.3 FlashX
被低估的早期信号:涨但无 X 热度
deeplethe/utopia(本地优先、agent 辅助的“文档→本体”知识工作台)和 sapientinc/PRAXIST(可执行的自主研究系统)在没有明显 X 带节奏的情况下自然上涨——通常意味着别人还没注意,alpha 在此,值得优先加入观察名单。
deeplethe/utopiasapientinc/PRAXIST
🎯 沿“agent 基础设施”主线深挖一层:对比 browser-use/jev-ultrafast(浏览器)与 trycua/cua(全 OS computer-use)两条路线,判断哪条更贴合你的用例;同时把无 X 热度却在涨的 deeplethe/utopia 加入重点观察。

📰 最新快讯

⭐ 多源共振

共振集中在“给 agent 用的工具”:浏览器(jev-ultrafast)、电脑操作(trycua/cua、arcbox)、技能包(security-audit-skill)、开发环境(coder/coder)、自主研究(PRAXIST)、知识工作台(utopia)。同一主线、多个团队同时发力。

trycua/cua ×5 ↺ 1d
githubxsocialboardhn
+273★/d 活跃开发 official #16
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
🔺 @trycua 首发 · 30h 前
Computer-use 2.0:跨 OS 驱动、agent 集群、评测与训练数据生成——“agent 基础设施”主线的核心一员,且为团队自荐(lead:trycua)。
githubxsocialboard
+985★/d 早期·低活动 official #1
🔺 @betterhn20 首发 · 16h 前
已核实=真实开源发布,非刷榜。Browser Use 出品的浏览器 agent,基于 TypeSafe“Jev/System One”模型(类型化概率决策、免视觉、70–500ms);自测 7s 完成 Google Flights、成本≈-90%(官方自测未复现)。新仓库 PR 少属正常。
githubxsocialboard
+550★/d 早期·低活动 official #17
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
🔺 @MaciejLukianski 首发 · 41h 前
Cloudflare 出品的 coding-agent 技能包:多阶段安全审计、独立验证、机器可读结论——“skill 即能力分发”的代表案例。
githubxsocialboard
+450★/d 活跃开发 official #3
The Photoshop alternative for Mac
🔺 @dotey 首发 · 5h 前
githubxsocialboard
+439★/d 活跃开发 official #2
deeplethe/utopia ×4 ↺ 1d
githubxsocialboard
+337★/d 活跃开发 official #9
Local-first, agent-assisted document-to-ontology workbench
本地优先、agent 辅助的文档→本体工作台;无 X 热度却涨=早期/被低估信号,优先加入观察。
hypit-ai/hypit ×4 ↺ 1d
githubxsociallaunch
+314★/d 活跃开发
🔺 @cccyd_qwq 首发 · 31h 前
已核实=真实项目(~1.3k★、活跃发版至 0.2.7)。让 coding agent 把短视频还原成可编辑“代码”、批量产出变体;以 `npx skills add` 技能包分发。工具+话题双热。
stablyai/orca ×4 ↺ 1d
githubxsocialboard
+267★/d 活跃开发 official #23
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
🔺 @GitTrend0x 首发 · 62h 前

🔥 动量榜

榜首多为刚开源的新仓库,PR/issue 少属正常现象(新项目),不等于刷榜。已逐一核实:jev-ultrafast、hypit 均为真实发布。

#Repo7d+1d★7d★质地官方X
1 browser-use/jev-ultrafast ↺ 1d
Python
🔺 @betterhn20 首发 · 16h 前
已核实=真实开源发布,非刷榜。Browser Use 出品的浏览器 agent,基于 TypeSafe“Jev/System One”模型(类型化概率决策、免视觉、70–500ms);自测 7s 完成 Google Flights、成本≈-90%(官方自测未复现)。新仓库 PR 少属正常。
+985 9,147 早期·低活动 #1 14×
2 cloudflare/security-audit-skill ↺ 1d
JavaScript · A coding-agent skill for multi-phase security audits with in
🔺 @MaciejLukianski 首发 · 41h 前
Cloudflare 出品的 coding-agent 技能包:多阶段安全审计、独立验证、机器可读结论——“skill 即能力分发”的代表案例。
+550 13,528 早期·低活动 #17 32×
3 robbietilton/Compositor ↺ 1d
Swift · The Photoshop alternative for Mac
🔺 @dotey 首发 · 5h 前
+450 2,495 活跃开发 #3
4 NandhaKishorM/laya ↺ 1d
Python
+439 1,664 活跃开发 #2
5 deepseek-ai/deepseek-harness ↺ 1d
TypeScript
🔺 @the_osps 首发 · 7h 前
+408 7,871 早期·低活动 10×
6 deeplethe/utopia ↺ 1d
Python · Local-first, agent-assisted document-to-ontology workbench
本地优先、agent 辅助的文档→本体工作台;无 X 热度却涨=早期/被低估信号,优先加入观察。
+337 2,474 活跃开发 #9 10×
7 hypit-ai/hypit ↺ 1d
TypeScript
🔺 @cccyd_qwq 首发 · 31h 前
已核实=真实项目(~1.3k★、活跃发版至 0.2.7)。让 coding agent 把短视频还原成可编辑“代码”、批量产出变体;以 `npx skills add` 技能包分发。工具+话题双热。
+314 10,935 活跃开发 28×
8 alibaba/open-code-review ↺ 1d
Go
🔺 @shao__meng 首发 · 71h 前
+302 14,364 活跃开发 22×
9 trycua/cua ↺ 1d
HTML · Scale computer-use 2.0 with open-source drivers, cross-OS fl
🔺 @trycua 首发 · 30h 前
Computer-use 2.0:跨 OS 驱动、agent 集群、评测与训练数据生成——“agent 基础设施”主线的核心一员,且为团队自荐(lead:trycua)。
+273 2,045 活跃开发 #16 20×
10 arcboxlabs/arcbox ↺ 1d
Rust · Run AI agents on real and isolated machines — own kernel, fi
+268 1,346 早期·低活动 #7
11 stablyai/orca ↺ 1d
TypeScript · Orca is the ADE for working with a fleet of parallel agents.
🔺 @GitTrend0x 首发 · 62h 前
+267 5,082 活跃开发 #23 21×
12 ruanyf/weekly 🆕 new
· 科技爱好者周刊,每周五发布
🔺 @clxymox 首发 · 7h 前
+261 931 活跃开发 #4
13 Open-Dev-Society/OpenStock ↺ 1d
TypeScript · OpenStock is an open-source alternative to expensive market
+239 2,013 早期·低活动 #20 12×
14 tt-a1i/archify ↺ 1d
JavaScript
🔺 @GitTrend0x 首发 · 62h 前
+230 6,975 活跃开发 12×
15 eternity4719/HowToLiveBetter ↺ 1d
HTML
🔺 @knowledgefxg 首发 · 60h 前
+221 5,936 早期·低活动 10×
16 TianyuCodings/NanoJev 🆕 new
Python · A nano replica of Jev: parallel decisions, dynamic candidate
🔺 @xx309212 首发 · 17h 前
+213 1,005 早期·低活动 #11
17 Tencent/WeKnora ↺ 1d
Go
🔺 @QingQ77 首发 · 20h 前
+196 4,752 活跃开发 17×
18 docling-project/docling 🆕 new
Python · Get your documents ready for gen AI
🔺 @TodayKan 首发 · 42h 前
+194 859 活跃开发 #13 11×
19 addyosmani/agent-skills ↺ 1d
JavaScript
🔺 @shanyanggm 首发 · 71h 前
+192 3,174 活跃开发 30×
20 vladelaina/BongoCat 🆕 new
C · 🩷 💘C × SDL3 × OpenGL, stir it up, mash it together! Bong~
+189 482 活跃开发 #8

🗞️ Hacker News

💬 V2EX

🐧 LINUX DO

👽 Reddit

🛠️ 技术源 · GitHub Trending/Lobsters

📥 AI 博客 · Newsletter

📢 电报精选

📚 科技周刊 新项目/工具自荐

🎯 Alpha 账号 X 上最早带火仓库的人

作者leadslead率仓库数
@shanyanggm160.6722
@xzbx88880.88
@shaw_stone7383270.76
@GitTrend0x70.78
@the_osps60.866
@xfubot60.58
@FrontieraTechIT50.717
@DataChaz514
@RepoGems50.834
@neil_xbt50.633
@key_indie50.832
@LFrefman50.368
@LoveAIbrain514
@iasg100450.3311
@bilawalsidhu411
@jasontopia40.83
@Sn0wbrave40.56
@shao__meng40.575
@vintcessun40.448
@seekjourney40.673
@ClaudeCodeLog411
@0x_Kratos311
@FreeYoung55202230.387
@JackAIStudio999311
@engmaxxing30.436

🏆 各领域最强模型

文本 / 对话
Claude Fable 5.1
Anthropic · 53.4
图像生成
GPT Image 2.5 Flare
OpenAI · 1188
图像编辑
GPT Image 2.5 Sunburst
OpenAI · 1176
文生视频
Wan 3.0
Alibaba · 1336
图生视频
Gemini Omni Flash
Google · 1369
语音合成
Sonic 3.6
Cartesia · 1276

🏆 能力排行榜 Artificial Analysis

文本 / 对话 Intelligence Index
  1. 1Claude Fable 5.153.4
  2. 2GPT-6 Astra52.7
  3. 3Claude Opus 550.8
  4. 4Claude Fable 549.6
  5. 5Muse Spark 1.348.1
  6. 6GPT-5.6 Sol47
  7. 7Qwen3.8 Max45.4
  8. 8GLM-5.344.8
  9. 9Grok 4.644.3
  10. 10Step 5 Preview43.7
  11. 11Kimi K343.6
  12. 12GPT-5.6 Terra42.1
图像生成 Text→Image Arena Elo
  1. 1GPT Image 2.5 Flare1188
  2. 2GPT Image 2.5 Sunburst1182
  3. 3GPT Image 21171
  4. 4Grok Imagine Image 2.01154
  5. 5MAI-Image-2.61147
  6. 6Reve 2.11129
  7. 7Nano Banana 21122
  8. 8Muse Image1111
  9. 9GPT Image 1.51102
  10. 10MAI-Image-2.51102
  11. 11Nano Banana Pro1100
  12. 12MAI-Image-2.6-Flash1099
图像编辑 Image-Editing Arena Elo
  1. 1GPT Image 2.5 Sunburst1176
  2. 2GPT Image 2.5 Flare1155
  3. 3MAI-Image-2.61132
  4. 4MAI-Image-2.6-Flash1122
  5. 5GPT Image 21121
  6. 6Muse Image1115
  7. 7MAI-Image-2.51113
  8. 8MAI-Image-2.5-Pro1106
  9. 9Seedream 5.0 Pro1106
  10. 10Nano Banana 21105
  11. 11Grok Imagine Image 2.01104
  12. 12GPT Image 1.51104
文生视频 Text→Video Arena Elo
  1. 1Wan 3.01336
  2. 2Gemini Omni Flash1330
  3. 3MiniMax H31302
  4. 4HappyHorse-1.01287
  5. 5HappyHorse-1.11272
  6. 6Dreamina Seedance 2.0 720p1259
  7. 7Wan2.7-2606121243
  8. 8grok-imagine-video1235
  9. 9Kling 3.0 Omni 1080p1230
  10. 10PixVerse V5.61230
  11. 11PixVerse V61230
  12. 12Kling 3.0 1080p1230
图生视频 Image→Video Arena Elo
  1. 1Gemini Omni Flash1369
  2. 2Wan 3.01361
  3. 3Bach 1.0 Pro1359
  4. 4MiniMax H31354
  5. 5PixVerse V61337
  6. 6Dreamina Seedance 2.0 720p1336
  7. 7grok-imagine-video-1.51329
  8. 8grok-imagine-video1326
  9. 9HappyHorse-1.11312
  10. 10Kling 2.5 Turbo 1080p1296
  11. 11HappyHorse-1.01293
  12. 12Vidu Q3 Pro1290
语音合成 Text→Speech Arena Elo
  1. 1Sonic 3.61276
  2. 2Qwen-Audio-3.0-TTS-Plus1260
  3. 3Realtime TTS-21247
  4. 4Simba 3.21240
  5. 5Luna TTS1231
  6. 6Realtime TTS-2 Flash1215
  7. 7StepAudio 2.5 TTS1209
  8. 8Breeze TTS 21205
  9. 9Gemini 3.1 Flash TTS1201
  10. 10v3 Conversational1197
  11. 11Sonic 3.51184
  12. 12Lightning V3.1 Pro1179

🧠 最新发布

2026-09-17
KAT-Coder-Pro V2.5
Kwaipilot
2026-09-17
Gemini Omni Flash Preview
Google
2026-09-17
Venice Uncensored
Venice
2026-09-17
Nano Banana 2 Lite
Google
2026-09-17
Jev 1.13
TypeSafe AI
2026-09-16
Union Alpha
Stealth
2026-09-12
Schematron V2 Turbo
Inference.net
2026-09-12
Schematron V2 Small
Inference.net

🧠 模型发布时间线

优先看 source_count≥2(多源确认)与 ★notable;留意“快变体”(如 GLM-5.3 FlashX)——呼应专用快模型分工趋势。

2026-09-17
KAT-Coder-Pro V2.5
Kwaipilot
aimlapi
2026-09-17
Gemini Omni Flash Preview
Google
aimlapi
2026-09-17
Venice Uncensored
Venice
aimlapi
2026-09-17
Nano Banana 2 Lite
Google
aimlapi
2026-09-17
Jev 1.13
TypeSafe AI
aimlapi
2026-09-16
Union Alpha
Stealth
aimlapi
2026-09-12
Schematron V2 Turbo
Inference.net
aimlapi
2026-09-12
Schematron V2 Small
Inference.net
aimlapi
2026-09-11
Fugu Ultra v2.0
Sakana AI
llmgateway
2026-09-11
Kimi K2.8 Preview
Moonshot AI
llmstats
2026-09-11
Atria Dawn Preview
Shanghai AI Laboratory
llmstatsllmgateway
2026-09-11
Fugu Ultra v2
Sakana AI
aimlapi
2026-09-11
Fugu Max
Sakana AI
aimlapillmgateway
2026-09-10 · ★
Ling 3.0 Flash VL
inclusionAI
aimlapiopper
2026-09-10 · ★
DeepSeek V4.1 Flash
DeepSeek AI
aimlapillmstatsopperllmgateway
2026-09-10
DeepSeek Chat (V4.1 Flash)
DeepSeek AI
aimlapi
2026-09-08
GPT Image 2.5 Sunburst
Open AI
aimlapillmgateway
2026-09-08
GPT Image 2.5 Flare
Open AI
aimlapillmgateway
2026-09-08
Mercury 2.5
Inception
aimlapi

📄 论文 PwC + arXiv

近期论文集中在 World Models / 开放式任务泛化 / 自主研究方向,与榜上“自主研究/computer-use”类项目相互印证。

🤗 HF 采用榜 下载/点赞

Modeldownloadslikespipeline
prism-ml/Ternary-Bonsai-2-27B-gguf1,516,9601,283text-generation
deepseek-ai/DeepSeek-V4.1-Flash482,2703,343image-text-to-text
convaiinnovations/laya0663text-classification
XingChen-AGI/Xing4.0-29B-A4B7,278708text-generation
Qwen/Qwen3.8-27B7,365,36815,787image-text-to-text
m-a-p/YuE2-3B15,446888text-to-audio
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF1,154,2651,438image-text-to-text
ukisai/Swift-Qwen3.8-27b8,761496image-text-to-text
harshatheg/Qwen-2.5-1B-RLCD0448text-generation
Lightricks/LTX-2.51,607,8154,468image-to-video
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF1,256,962979image-text-to-text
unsloth/Qwen3.8-27B-GGUF7,118,3634,394
openbmb/MiniCPM5-2B389,5551,603text-generation
ukisai/Swift-Qwen3.8-27B-GGUF120,740320image-text-to-text
Qwen/Qwen3.8-Flash-Next742,5865,463image-text-to-text
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit23,111261text-generation
meta-llama/Llama-3.1-8B-Instruct5,919,7467,756text-generation
AlexWortega/openjev0236text-classification
Edge0/Edge0-35B-A3B-preview68,4033,509text-generation
dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP834,230313image-text-to-text
MiniMaxAI/MiniMax-H34,299,7375,501image-text-to-video
TokenRhythm/NeoHorse-1-9B11,692971text-generation
internlm/Atria-Dawn-Preview806204
TaichuAI/ZDTaichu5.0-9B2,926206image-text-to-text
tencent/AuK3,355334text-to-speech
WarmBloodAban/Minimax-h3_Singularity231,197545image-to-video
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF31,099177image-text-to-text
zai-org/GLM-5.3-Flash2,905,9322,484image-text-to-text
sentence-transformers/all-MiniLM-L6-v2254,149,2356,090sentence-similarity
Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v40156

⚡ System One 决策模型 314 项目 · Jev/TypeSafe 生态 · 快决策(非推理)

SDK & Decision Frameworks 56Evaluation & Observability 29Browser & OS Action 26High-Frequency & Simulation 26Routing & Cost Optimization 23Security & Guardrails 22Context GC & Filter 19Domain & Vertical Tools 19MCP & Integrations 18CLI & Pipelines 15Data & Search 15Decision Tools 12Codebase & Graph Pathfinding 11Creative Tools 11SDK & Integrations 6Voice & Conversation 4Classification & Taxonomy 2
项目类别Jev 决策点
langchain langchain-ai146,634SDK & IntegrationsSubmits binary, categorical and ordered-score questions and returns typed answers with probabilities.
ai-hedge-fund virattt63,505Domain & Vertical ToolsConverts strategy questions to System One requests and normalizes native answers to the project’s result format.
litellm BerriAI59,123Routing & Cost OptimizationMaps requests to configured complexity classes that drive backend routing.
oh-my-pi can135731,850Routing & Cost OptimizationSends agent state and typed questions to Jev and parses structured answers.
jev-model-router davila730,779Routing & Cost OptimizationEvaluates task tier, reasoning needs and production risk; local policy maps results to invocation settings.
composio ComposioHQ30,238SDK & Decision FrameworksTurns tool or action conditions into structured questions and passes Jev answers to local invocation logic.
ai vercel26,835SDK & Decision FrameworksMaps choice, score, and yes/no questions to TypeSafe System One requests and parses typed results.
cua trycua23,683Browser & OS ActionReads DOM or supported visual-region descriptions and returns a supplied candidate action ID.
pydantic-ai pydantic20,035SDK & IntegrationsConverts supported structured output fields into typed Jev questions and maps answers back to the output model.
eliza elizaOS19,361SDK & Decision FrameworksOnly an explicit systemOne call sends state and questions, returning validated typed answers.
langchainjs langchain-ai18,210SDK & Decision FrameworksUses invoke to call TypeSafe and parse choice, noul, score and probability fields.
json-render vercel-labs16,572Creative ToolsEvaluates component configurations through Vercel AI Gateway, then composeSpec assembles the UI specification.
openchamber openchamber10,060Routing & Cost OptimizationJev selects a task category; local category mappings determine the model configuration.
rig-typesafeai 0xPlaygrounds8,669SDK & Decision FrameworksSends application state and questions to Jev and parses Choice, Score or Noul answers.
firstmate kunchenguid6,587Routing & Cost OptimizationSends the task brief and candidate rules to Jev, then resolves execution profiles with confidence and local conditions.
jev-ultrafast browser-use6,031Browser & OS ActionChooses an action and its matching DOM element in one request; a text model generates input text.
agentgateway agentgateway4,926Security & GuardrailsJev scores jailbreaks, harmful content and secret disclosure; thresholds or evaluation errors reject requests.
latitude-llm latitude-dev4,655Evaluation & ObservabilityJudges which checks apply and can add checks when thresholds and rate limits permit.
fast-jev-compaction tamaratran3,528Context GC & FilterSeparately judges whether a tool call and its full output are still needed; code keeps, truncates or drops them.
ax ax-llm2,926SDK & IntegrationsMaps supported signatures to Jev questions or sends native System One requests.
vellum-assistant vellum-ai1,287MCP & IntegrationsSubmits state and question bundles to System One and returns structured answers to the Assistant.
jev-desktop lahfir1,275Browser & OS ActionJev selects a target and action and estimates presence and risk; local policy decides whether to execute.
celesto CelestoAI943Codebase & Graph PathfindingJudges whether a finding was introduced by the change, is supported and merits a fix.
jev-trader jarrodwatts936Domain & Vertical ToolsIn Jev mode, order-book judgments feed code that simulates fills or submits configured post-only limit orders.
atomic bastani-inc806Routing & Cost OptimizationSends predefined questions to Jev and decodes answers for callers; regular models still generate code.
aiavatarkit uezo676Voice & ConversationAssesses utterance completeness and whether the user is likely to continue speaking.
NanoJev TianyuCodings657High-Frequency & SimulationEvaluates multiple questions and dynamic candidate spaces concurrently in a single forward pass, logging navigation choices.
kody kentcdodds654Data & SearchSends a Score question per candidate, reorders and drops low scores; the model id is typesafe/jev.
Agent AgentiLoop616Security & GuardrailsAdds a destructive-risk judgment after local shell checks and refuses commands above the configured threshold.
req_llm agentjido577SDK & Decision FrameworksSends state and questions, normalizes answers and retains the raw provider response.
omg.dev BennyKok531Browser & OS ActionChooses controls and checks completion or blockage before the test runner operates the UI.
vexjoy-agent notque420Routing & Cost OptimizationAfter deterministic routing guards, Jev judges the remaining candidates and required workflow components.
foreman thruwire344CLI & PipelinesAsyncTypeSafeClient.system_one with default jev-latest sends Noul questions for supervision.
WrongStack WrongStack329Routing & Cost OptimizationJev evaluates the task against eligible specialists; local dispatch rules use the result.
instructor-php cognesy327SDK & Decision FrameworksConverts application state and typed questions into Jev requests and maps responses to PHP decision objects.
kev jaredpalmer310High-Frequency & SimulationAttaches a parallel decision head to an open 0.5B model to answer discrete questions directly from token activations.
typesafe-computer-use awlevin302Browser & OS ActionSelects the next step from deterministically extracted controls and actions before desktop execution.
Jev Review devagrawal09284Codebase & Graph PathfindingJudges risk, files, evidence regions, mechanisms and severity before rule-based reviewer routing.
orchestkit yonatangross278CLI & PipelinesClassifies the first task prompt and branch state by work type; local policy accepts the result or falls back.
typesafe-mario fhshaik266High-Frequency & SimulationReads motion, enemies, terrain and recent controls, then selects a predefined legal action.

🅱️ B站 AI 无限竞技场 18 模型 · 夺冠率

#模型夺冠率冠/测
1GPT-6 Astra OpenAI55%12/22
2Claude Fable 5.1 Anthropic50%6/12
3GLM-5.3 Z.ai16%3/19
3GPT-5.6 Sol OpenAI16%3/19
5Claude Fable 5 Anthropic27%3/11
6Kimi K3 Moonshot10%2/21
7Claude Opus 5 Anthropic11%2/19
8DeepSeek-V4-Flash DeepSeek13%2/16
9DeepSeek-V4-Pro DeepSeek5%1/21
10Qwen3.8-Max Alibaba6%1/18
11DeepSeek V4.1 Flash DeepSeek8%1/13
12Gemini 3.8 Flash Google8%1/12
13Hy 4 Tencent14%1/7
14Gemini 3.7 Flash Google20%1/5
15GPT-5.6 Terra OpenAI25%1/4
15Seed-2.0 pro ByteDance25%1/4
17Doubao-Seed-Evolving ByteDance33%1/3
18Seed-2.0 Mini ByteDance100%1/1
👤 AI UP主 从赛题发现,点击直达主页
🎯 赛题

🎬 AI 视频 订阅 · 搜索 · B站,分开组织

📌 订阅频道 4 个博主 · 最新上传
📺 Best Partners TV 12
📺 AI超元域 12
🔎 搜索发现 相关度+播放量筛选 · 非订阅
🅱️ B站 AI 竞技场 按播放量

🚀 产品发布 whatships · What's Launch

📦 版本发布 tracked repos releases

🛰️ Skywork 动态

🧪 Show HN 开发者发布的新产品

💰 商业动态 · TechCrunch/VB/MIT TR