⚡ GitHub Pulse

生成于 2026-09-27 17:17 UTC · 追踪 5,028 仓库 · 8 多源共振
🔍 搜索中 · 显示所有标签页的匹配项 · 按 Esc 清除
📊 深度分析(2026-09-20 20:40 北京)基于更早的数据快照,数据已刷新,已自动隐藏过时结论 —— 待重新生成分析后显示。

💎 高价值精选 AI 判断 · 跨源挖掘

📰 最新快讯

⭐ 多源共振

githubxsocialboardlaunch
+1,798★/d 活跃开发 official #9
The open-source app everyone uses to manage agents at work
🔺 @shaw_stone73832 首发 · 9h 前
dream-num/univer ×5 ↺ 4d
githubxsocialboardlaunch
+801★/d 活跃开发 official #7
The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
🔺 @tylerrwayne 首发 · 56h 前
reladraw/reladraw ×5 🆕 new
githubxsocialboardhn
+468★/d 早期·低活动 official #8
🔺 @betterhn20 首发 · 1h 前
githubxsocialboard
+4,030★/d 活跃开发 official #2
Hindsight: Agent Memory That Learns
🔺 @engmaxxing 首发 · 56h 前
githubxsocialboard
+2,475★/d 活跃开发 official #3
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
🔺 @AI_Jasonyu 首发 · 23h 前
githubxsocialboard
+2,110★/d 早期·低活动 official #23
354 条循证建议,覆盖长寿防病、急救、省钱理财、法律红线、失业与工伤、医保社保、恋爱婚育、出国与技能。每条写明成本、收益、证据等级和原始出处,只引期刊论文与官方文件。
🔺 @Coooperni 首发 · 64h 前
githubxsocialboard
+554★/d 活跃开发 official #10
Continual learning infra for self-improving agents
🔺 @kimmonismus 首发 · 54h 前
githubxsocialboard
+553★/d 早期·低活动 official #25
Suitable for Android APK reverse engineering analysis
🔺 @weiwei2018831 首发 · 20h 前

🔥 动量榜

#Repo7d+1d★7d★质地官方X
1 vectorize-io/hindsight ↺ 4d
Python · Hindsight: Agent Memory That Learns
🔺 @engmaxxing 首发 · 56h 前
+4,030 12,082 活跃开发 #2 14×
2 debpalash/VoiceStudio 🆕 new
Python · VoiceStudio is the open-source, fully-local ElevenLabs alter
🔺 @AI_Jasonyu 首发 · 23h 前
+2,475 5,690 活跃开发 #3 41×
3 hydra-db/hydradb ↺ 4d
Rust · HydraDB - fast graph database on object storage
🔺 @iasg1004 首发 · 70h 前
+2,470 7,270 早期·低活动 #5 2×
4 rocketride-org/rocketride-server ↺ 4d
Python · High-performance AI pipeline engine with a C++ core and 50+
+2,366 6,942 活跃开发 #4 1×
5 eternity4719/HowToLiveBetter ↺ 8d
HTML · 354 条循证建议,覆盖长寿防病、急救、省钱理财、法律红线、失业与工伤、医保社保、恋爱婚育、出国与技能。每条写明成本、收
🔺 @Coooperni 首发 · 64h 前
+2,110 11,805 早期·低活动 #23 41×
6 paperclipai/paperclip ↺ 3d
TypeScript · The open-source app everyone uses to manage agents at work
🔺 @shaw_stone73832 首发 · 9h 前
+1,798 7,884 活跃开发 #9 22×
7 cdyforever/how-to-live-better 🆕 new
HTML · 《高性价比人生指南》全书 528 条的在线单页阅读版:手机可读、可搜索、零依赖、支持离线
🔺 @aigeeknews 首发 · 39h 前
+1,026 2,011 早期·低活动 #14 3×
8 dream-num/univer ↺ 4d
TypeScript · The Office Harness for AI Agents — Spreadsheets, Docs, Slide
🔺 @tylerrwayne 首发 · 56h 前
+801 5,489 活跃开发 #7 15×
9 latent-spaces/brag ↺ 8d
Python
🔺 @itsvlady 首发 · 8h 前
+793 4,402 早期·低活动 – 14×
10 deepseek-ai/deepseek-harness ↺ 8d
TypeScript
+713 6,364 早期·低活动 – 19×
11 mexicat/pdoom-video 🆕 new
TypeScript
🔺 @OmNawale45831 首发 · 9h 前
+689 728 早期·低活动 – 5×
12 NandhaKishorM/laya ↺ 8d
Python
🔺 @clxymox 首发 · 61h 前
+667 22,520 活跃开发 – 36×
13 rohitg00/ai-engineering-from-scratch ↺ 3d
Python
🔺 @swamiabhishek45 首发 · 41h 前
+627 3,876 活跃开发 – 15×
14 juspay/hyperswitch ↺ 2d
Rust · Open source, composable payments platform | PCI compliant |
+566 1,180 活跃开发 #19 5×
15 Human-Agent-Society/reef ↺ 2d
Python · Continual learning infra for self-improving agents
🔺 @kimmonismus 首发 · 54h 前
+554 2,546 活跃开发 #10 10×
16 newliver666/apk-reverse ↺ 4d
Python · Suitable for Android APK reverse engineering analysis
🔺 @weiwei2018831 首发 · 20h 前
+553 2,132 早期·低活动 #25 1×
17 pacifio/atlas ↺ 6d
Rust
+531 3,203 活跃开发 – 10×
18 stablyai/orca ↺ 8d
TypeScript
🔺 @shanyanggm 首发 · 66h 前
+478 5,829 活跃开发 – 25×
19 reladraw/reladraw 🆕 new
TypeScript
🔺 @betterhn20 首发 · 1h 前
+468 687 早期·低活动 #8 2×
20 mobile-next/mobile-mcp 🆕 new
TypeScript · Model Context Protocol Server for Mobile Automation and Scra
+424 975 活跃开发 #18 10×

🗞️ Hacker News

🛠️ 技术源 · GitHub Trending/Lobsters

🐧 LINUX DO

👽 Reddit

💬 V2EX

📥 AI 博客 · Newsletter

📢 电报精选

📚 科技周刊 新项目/工具自荐

🎯 Alpha 账号 X 上最早带火仓库的人

作者leadslead率仓库数
@shanyanggm390.7527
@iasg1004230.4727
@shaw_stone73832170.5913
@xzbx888150.888
@xfubot140.514
@the_rza_140.4813
@GitTrend0x130.769
@QingQ77130.4316
@engmaxxing120.610
@vintcessun120.3622
@the_osps110.6110
@Sn0wbrave100.598
@clxymox100.510
@aigeeknews100.3716
@shao__meng100.677
@newlinedotco90.645
@jasontopia80.893
@FreeYoung55202280.537
@LFrefman80.479
@FrontieraTechIT70.648
@DataChaz70.78
@betterhn5070.586
@betterhn2070.479
@Angle1991Ai70.73
@seekjourney70.74

🏆 各领域最强模型

文本 / 对话
Claude Opus 5.5
Anthropic · 57.6
图像生成
GPT Image 2.5 Sunburst
OpenAI · 1196
图像编辑
GPT Image 2.5 Sunburst
OpenAI · 1180
文生视频
Wan 3.0
Alibaba · 1335
图生视频
Gemini Omni Flash
Google · 1369
语音合成
Sonic 3.6
Cartesia · 1277

🏆 能力排行榜 Artificial Analysis

文本 / 对话 Intelligence Index
  1. 1Claude Opus 5.557.6
  2. 2Claude Fable 5.153.4
  3. 3GPT-6 Astra52.7
  4. 4Claude Opus 550.8
  5. 5Claude Fable 549.6
  6. 6Muse Spark 1.348.1
  7. 7GPT-6 Sol47.5
  8. 8GPT-5.6 Sol47
  9. 9Grok 4.746.4
  10. 10MiMo-V2.6-Pro46.3
  11. 11Qwen3.8 Max45.4
  12. 12GLM-5.344.8
图像生成 Text→Image Arena Elo
  1. 1GPT Image 2.5 Sunburst1196
  2. 2GPT Image 2.5 Flare1190
  3. 3GPT Image 21171
  4. 4Grok Imagine Image 2.01157
  5. 5MAI-Image-2.61148
  6. 6Nano Banana 21123
  7. 7Muse Image1111
  8. 8GPT Image 1.51104
  9. 9MAI-Image-2.51102
  10. 10MAI-Image-2.6-Flash1102
  11. 11Nano Banana Pro1101
  12. 12MAI-Image-2.5-Pro1100
图像编辑 Image-Editing Arena Elo
  1. 1GPT Image 2.5 Sunburst1180
  2. 2GPT Image 2.5 Flare1161
  3. 3MAI-Image-2.61134
  4. 4MAI-Image-2.6-Flash1124
  5. 5GPT Image 21121
  6. 6Muse Image1117
  7. 7MAI-Image-2.51112
  8. 8Nano Banana 21107
  9. 9Seedream 5.0 Pro1106
  10. 10MAI-Image-2.5-Pro1105
  11. 11Grok Imagine Image 2.01105
  12. 12GPT Image 1.51102
文生视频 Text→Video Arena Elo
  1. 1Wan 3.01335
  2. 2Gemini Omni Flash1332
  3. 3MiniMax H31302
  4. 4HappyHorse-1.01286
  5. 5HappyHorse-1.11272
  6. 6Dreamina Seedance 2.0 720p1259
  7. 7Wan2.7-2606121243
  8. 8grok-imagine-video1235
  9. 9PixVerse V5.61230
  10. 10PixVerse V61230
  11. 11Kling 3.0 Omni 1080p1230
  12. 12Kling 3.0 1080p1230
图生视频 Image→Video Arena Elo
  1. 1Gemini Omni Flash1369
  2. 2Bach 1.0 Pro1362
  3. 3Wan 3.01362
  4. 4MiniMax H31358
  5. 5PixVerse V61340
  6. 6Dreamina Seedance 2.0 720p1340
  7. 7grok-imagine-video-1.51330
  8. 8grok-imagine-video1329
  9. 9HappyHorse-1.11314
  10. 10Kling 2.5 Turbo 1080p1299
  11. 11HappyHorse-1.01295
  12. 12Vidu Q3 Pro1291
语音合成 Text→Speech Arena Elo
  1. 1Sonic 3.61277
  2. 2Gemini 3.8 Flash TTS1268
  3. 3Qwen-Audio-3.0-TTS-Plus1259
  4. 4Realtime TTS-21244
  5. 5Gemini 3.8 Flash-Lite TTS1240
  6. 6Simba 3.21238
  7. 7Luna TTS1232
  8. 8Realtime TTS-2 Flash1212
  9. 9Breeze TTS 21206
  10. 10Gemini 3.1 Flash TTS1202
  11. 11StepAudio 2.5 TTS1200
  12. 12v3 Conversational1197

🥇 综合能力榜 Benchmark Heaven · 7 榜合一(AA+Epoch ECI+DesignArena)

#模型综合分$/1MBenchmaxxing
1Claude Opus 5.5 Anthropic 🏅前沿100$5.45—
2Claude Fable 5.1 Anthropic97.2$13.640.58
3GPT 6 Astra OpenAI97.2$13.64-7.58
4Claude Opus 5 Anthropic94.8$6.82-5.24
5Claude Fable 5 Anthropic94.2$13.64-4.41
6Muse Spark 1.3 Meta86.1$1.525.71
7GPT 6 Sol OpenAI89.2$2.73—
8GPT 5.6 Sol OpenAI 🏅前沿93.7$2.73-2.41
9Grok 4.7 xAI89.1$1.89—
10MiMo V2.6 Pro Xiaomi 开源 🏅前沿91.7$0.47—
11Qwen3.8 Max 0902 Alibaba86$2.36—
12GLM 5.3 Z.ai 开源 🏅前沿84.6$0.45-0.08
13Grok 4.6 xAI82.9$2.365.38
14Step 5 Preview StepFun84.5$1.15—
15Kimi K3 Moonshot AI 开源87.7$1.733.68

Benchmaxxing:BH 的指标,+ 值越高表示该模型在"公开基准"上比"不可训练的封闭题"排名更靠前(BH 定义与计算,非本站判断)。

💰 性价比 / 最省钱 跨 provider 最低价 · 10:1 blended · 综合分≥60

模型$/1M综合分最便宜 provider
GLM 5.3 Flash 🏅前沿 🔒不训练 开源$0.0576.8InferenceNet
DeepSeek V4 Flash 0731 🔒不训练 开源$0.0572.8Sail Research
DeepSeek V4.1 Flash 🔒不训练 开源$0.0671.7InferenceNet
GPT 6 Luna 🏅前沿$0.1477.4OpenAI
Qwen3.8 Flash Next 🏅前沿 开源$0.1880.6Alibaba
Qwen3.8 27B 开源$0.2667.5Darkbloom
GPT 5.6 Luna 🔒不训练$0.2977.1Azure AI Foundry
DeepSeek V4 Pro 0813 开源$0.2974.9Baidu
DeepSeek V4 Pro 开源$0.3861.5Baidu
Solar Pro 4$0.3862
GLM 5.3 🏅前沿 开源$0.4584.6Baidu
MiMo V2.6 Pro 🏅前沿 开源$0.4791.7GMICloud
DeepSeek V4 Flash Vision$0.5264.2
Inkling Small 🔒不训练 开源$0.5263.4DeepInfra
Apodex 1.1$0.5565

跨 94 家 provider(含 21 家 🇪🇺 EU、54 家 🔒不训练)· 数据 2026-09-27

🏅 性价比前沿 没有更便宜的模型能在能力上胜过它们(Pareto)

  • Ling 3.0 Flash 综合 53.4 · $0.02/M · Novita · 开源
  • GLM 5.3 Flash 综合 76.8 · $0.05/M · InferenceNet · 开源
  • GPT 6 Luna 综合 77.4 · $0.14/M · OpenAI
  • Qwen3.8 Flash Next 综合 80.6 · $0.18/M · Alibaba · 开源
  • GLM 5.3 综合 84.6 · $0.45/M · Baidu · 开源
  • MiMo V2.6 Pro 综合 91.7 · $0.47/M · GMICloud · 开源
  • GPT 5.6 Sol 综合 93.7 · $2.73/M · OpenAI
  • Claude Opus 5.5 综合 100 · $5.45/M · Anthropic

🏢 官方发布 · 第一手 各厂官网直发 · 按时间融合 · 每家限量防刷屏

🧠 最新发布 跨源去重时间线

2026-09-25
Smart Route
LLM Gateway
2026-09-23
Aion-3.5-Mini
AionLabs
2026-09-23
Ember-1
Fireworks
2026-09-23
Qwen3.8 Max Prime
Alibaba Cloud
2026-09-23
Aion-3.5
AionLabs
2026-09-23
Space Bunny Alpha
Stealth
2026-09-22
Solar Mini 4
Upstage
2026-09-22
GPT-6 Sol
Open AI

🧠 模型发布时间线

2026-09-25
Smart Route
LLM Gateway
llmgateway
2026-09-23
Aion-3.5-Mini
AionLabs
aimlapi
2026-09-23
Ember-1
Fireworks
aimlapi
2026-09-23
Qwen3.8 Max Prime
Alibaba Cloud
aimlapi
2026-09-23
Aion-3.5
AionLabs
aimlapi
2026-09-23
Space Bunny Alpha
Stealth
aimlapi
2026-09-22
Solar Mini 4
Upstage
aimlapi
2026-09-22 · ★
GPT-6 Sol
Open AI
aimlapillmstatsopperllmgateway4×
2026-09-22 · ★
GPT-6 Luna
Open AI
aimlapillmstatsopperllmgateway4×
2026-09-21 · ★
Claude Opus 5.5
Anthropic
aimlapillmstatsopperllmgateway4×
2026-09-21
MiMo V2.6 Pro UltraSpeed
Xiaomi
aimlapi
2026-09-21 · ★
MiMo V2.6 Pro
Xiaomi
aimlapillmstatsopperllmgateway4×
2026-09-21 · ★
MiMo V2.6 Flash
Xiaomi
aimlapillmstatsopperllmgateway4×
2026-09-18
GLM 5.3 FlashX
Zhipu AI
aimlapi
2026-09-18
Ternary Bonsai 2 27B
PrismML
aimlapi
2026-09-17
KAT-Coder-Pro V2.5
Kwaipilot
aimlapi
2026-09-17
Gemini Omni Flash Preview
Google
aimlapi
2026-09-17
Venice Uncensored
Venice
aimlapi
2026-09-17
Nano Banana 2 Lite
Google
aimlapi
2026-09-17
Pareto
Unbiased
aimlapi
2026-09-17
Qwen3.8 Omni Flash
Alibaba Cloud
aimlapi
2026-09-16
Union Alpha
Stealth
aimlapi
2026-09-15
Jev 1.13
TypeSafe AI
aimlapillmgateway2×
2026-09-14 · ★
Grok 4.7
X AI
aimlapillmstatsopperllmgateway4×

📄 论文 PwC + arXiv

🤗 HF 采用榜 下载/点赞

Modeldownloadslikespipeline
convaiinnovations/laya04,045text-classification
Qwen/Qwen-Image-2.152,8042,466text-to-image
abenzerps/Qwen-Image-2.1-Uncensored-GGUF964,2202,036text-to-image
XingChen-AGI/Xing4.0-29B-A4B45,0281,776text-generation
abenzerps/Qwen-Image-2.1-GGUF182,313988text-to-image
Edge0/Audio8-ASR-Infinite19,434973automatic-speech-recognition
Altworld/Hemmingway-15,904727text-generation
prism-ml/Ternary-Bonsai-2-27B-gguf3,343,7482,173text-generation
Comfy-Org/Qwen-Image-2.13,987,373796
XiaomiMiMo/MiMo-V2.6-Pro-RL75,079545text-generation
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B8,839518image-text-to-text
XiaomiMiMo/MiMo-V2.6-Flash-RL25,661489text-generation
TaichuAI/ZDTaichu5.0-9B11,6121,677image-text-to-text
StarDoc-AI/TeleOCR27,837550image-text-to-text
Qwen/Qwen3.8-27B6,727,62916,401image-text-to-text
nvidia/Nemotron-3-Diarization22,514393voice-activity-detection
Contrastive-LM/CLM-v0.1-8B766382text-ranking
Lightricks/LTX-2.51,601,0895,302image-to-video
Viggle/Qwen-Image-2.1-viggle-turbo133,151327text-to-image
AlexWortega/openjev0607text-classification
inclusionAI/Ming-Image-0.1-Design0292text-to-image
netease-youdao/Confucius4-R2T28,243432automatic-speech-recognition
pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF145,246292
yandex/AliceAI-Foundation-80B-A3B-Base3,456342text-generation
deepseek-ai/DeepSeek-V4.1-Flash651,0783,800image-text-to-text
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF1,608,4391,756image-text-to-text
akhilaaa3/Jev-Omni248266text-classification
unsloth/Qwen-Image-2.1-GGUF194,341265text-to-image
convaiinnovations/laya-multilingual0301text-classification
apple/LensVLM-9B1,740238image-text-to-text

⚡ System One 决策模型 695 项目 · Jev/TypeSafe 生态 · 快决策(非推理)

SDK & Decision Frameworks 121Domain & Vertical Tools 69CLI & Pipelines 62Security & Guardrails 58Routing & Cost Optimization 55High-Frequency & Simulation 52Browser & OS Action 50Data & Search 45MCP & Integrations 43Context GC & Filter 40Evaluation & Observability 29Decision Tools 25Creative Tools 20Codebase & Graph Pathfinding 14SDK & Integrations 6Voice & Conversation 4Classification & Taxonomy 2
项目★类别Jev 决策点
langchain langchain-ai146,956SDK & IntegrationsSubmits binary, categorical and ordered-score questions and returns typed answers with probabilities.
ai-hedge-fund virattt63,699Domain & Vertical ToolsConverts strategy questions to System One requests and normalizes native answers to the project’s result format.
litellm BerriAI59,511Routing & Cost OptimizationMaps requests to configured complexity classes that drive backend routing.
oh-my-pi can135733,041Routing & Cost OptimizationSends agent state and typed questions to Jev and parses structured answers.
jev-model-router davila731,561Routing & Cost OptimizationEvaluates task tier, reasoning needs and production risk; local policy maps results to invocation settings.
composio ComposioHQ30,300SDK & Decision FrameworksTurns tool or action conditions into structured questions and passes Jev answers to local invocation logic.
ai vercel26,922SDK & Decision FrameworksMaps choice, score, and yes/no questions to TypeSafe System One requests and parses typed results.
cua trycua26,156Browser & OS ActionReads DOM or supported visual-region descriptions and returns a supplied candidate action ID.
pydantic-ai pydantic20,133SDK & IntegrationsConverts supported structured output fields into typed Jev questions and maps answers back to the output model.
eliza elizaOS19,468SDK & Decision FrameworksOnly an explicit systemOne call sends state and questions, returning validated typed answers.
jev-ultrafast browser-use19,223Browser & OS ActionChooses an action and its matching DOM element in one request; a text model generates input text.
langchainjs langchain-ai18,223SDK & Decision FrameworksUses invoke to call TypeSafe and parse choice, noul, score and probability fields.
json-render vercel-labs18,200Creative ToolsEvaluates component configurations through Vercel AI Gateway, then composeSpec assembles the UI specification.
openchamber openchamber10,398Routing & Cost OptimizationJev selects a task category; local category mappings determine the model configuration.
rig-typesafeai 0xPlaygrounds8,713SDK & Decision FrameworksSends application state and questions to Jev and parses Choice, Score or Noul answers.
deep-searcher zilliztech8,282Data & SearchJev returns a structured decision for the local program; consult the source for the exact decision policy.
GPTCache zilliztech8,201Data & SearchJev returns a structured decision for the local program; consult the source for the exact decision policy.
firstmate kunchenguid7,079Routing & Cost OptimizationSends the task brief and candidate rules to Jev, then resolves execution profiles with confidence and local conditions.
fast-jev-compaction tamaratran6,607Context GC & FilterSeparately judges whether a tool call and its full output are still needed; code keeps, truncates or drops them.
kev jaredpalmer6,183High-Frequency & SimulationAttaches a parallel decision head to an open 0.5B model to answer discrete questions directly from token activations.
agentgateway agentgateway5,015Security & GuardrailsJev scores jailbreaks, harmful content and secret disclosure; thresholds or evaluation errors reject requests.
latitude-llm latitude-dev4,672Evaluation & ObservabilityJudges which checks apply and can add checks when thresholds and rate limits permit.
ax ax-llm2,946SDK & IntegrationsMaps supported signatures to Jev questions or sends native System One requests.
memsearch zilliztech2,649Data & SearchJev returns a structured decision for the local program; consult the source for the exact decision policy.
bootcamp milvus-io2,445Data & SearchJev returns a structured decision for the local program; consult the source for the exact decision policy.
jev-trader jarrodwatts2,207Domain & Vertical ToolsIn Jev mode, order-book judgments feed code that simulates fills or submits configured post-only limit orders.
NanoJev TianyuCodings2,135High-Frequency & SimulationEvaluates multiple questions and dynamic candidate spaces concurrently in a single forward pass, logging navigation choices.
jev-desktop lahfir1,604Browser & OS ActionJev selects a target and action and estimates presence and risk; local policy decides whether to execute.
vellum-assistant vellum-ai1,313MCP & IntegrationsSubmits state and question bundles to System One and returns structured answers to the Assistant.
celesto CelestoAI971Codebase & Graph PathfindingJudges whether a finding was introduced by the change, is supported and merits a fix.
typesafe-computer-use awlevin930Browser & OS ActionSelects the next step from deterministically extracted controls and actions before desktop execution.
jev feder-cr879SDK & Decision FrameworksJev returns a structured decision for the local program; consult the source for the exact decision policy.
atomic bastani-inc820Routing & Cost OptimizationSends predefined questions to Jev and decodes answers for callers; regular models still generate code.
hermes-jev-skills kerpopule716Routing & Cost OptimizationJev returns a structured decision for the local program; consult the source for the exact decision policy.
aiavatarkit uezo678Voice & ConversationAssesses utterance completeness and whether the user is likely to continue speaking.
kody kentcdodds672Data & SearchSends a Score question per candidate, reorders and drops low scores; the model id is typesafe/jev.
Agent AgentiLoop632Security & GuardrailsAdds a destructive-risk judgment after local shell checks and refuses commands above the configured threshold.
Jev-cu Sac-Y589Browser & OS ActionChooses targets and actions and assesses completion and risk; local policy controls execution or confirmation.
Jev Review devagrawal09582Codebase & Graph PathfindingJudges risk, files, evidence regions, mechanisms and severity before rule-based reviewer routing.
req_llm agentjido582SDK & Decision FrameworksSends state and questions, normalizes answers and retains the raw provider response.
🏅 JevBench 决策榜 JevBench v1 · 242 决策/模型 · 私有留出集 · 准确率 + 校准(Brier↓ 越低越准)
#系统准确率Brier↓类型
1Jev 1.13.0 (TypeSafe AI) TypeSafe AI 开源100%0.0028jev
2openjev-sglang (Qwen3.6-35B-A3B on SGLang) ekzhang 开源100%0.0305jev-rebuild
3system-one-open (Gemma 4 E2B LoRA on an L4) mithalouni 开源96%0.0851jev-rebuild
4open-alternative-jev (Qwen3.5-4B, HF Space) IkerMoel 开源——jev-rebuild
5open-jev-deberta-v3-large (local CPU) Kotoba Labs 开源67%0.4769jev-rebuild
6GPT-5.6 Luna (low reasoning effort) OpenAI 开源100%0.0003llm-baseline
7Gemini 3.1 Flash-Lite Google 开源100%0.0017llm-baseline
8DeepSeek V4.1 Flash (thinking default) DeepSeek 开源100%0.0001llm-baseline
9Qwen3.8 27B (Chutes TEE) Qwen / Chutes 开源100%0.0001llm-baseline

JevBench(BH):把"决策模型"(Jev/System One)与通用 LLM 放在同一批决策任务上比准确率与概率校准。

🅱️ B站 AI 无限竞技场 20 模型 · 夺冠率

#模型夺冠率冠/测
1GPT-6 Astra OpenAI50%16/32
2Claude Opus 5.5 Anthropic82%9/11
3Claude Fable 5.1 Anthropic21%4/19
4GLM-5.3 Z.ai12%3/26
5Claude Opus 5 Anthropic13%3/23
5GPT-5.6 Sol OpenAI13%3/23
7DeepSeek V4.1 Flash DeepSeek14%3/22
8Kimi K3 Moonshot7%2/27
9Claude Fable 5 Anthropic18%2/11
10DeepSeek-V4-Pro DeepSeek4%1/28
11Qwen3.8-Max Alibaba5%1/22
12DeepSeek-V4-Flash DeepSeek5%1/19
13Gemini 3.8 Flash Google6%1/18
14Hy 3 Tencent14%1/7
14Hy 4 Tencent14%1/7
16GPT-5.6 Terra OpenAI17%1/6
16Qwen3.7-Max Alibaba17%1/6
18Gemini 3.7 Flash Google25%1/4
19Seed-2.0 pro 33%1/3
20Doubao-Seed-Evolving 50%1/2
👤 AI UP主 从赛题发现,点击直达主页
🎯 赛题

🎬 AI 视频 订阅 · 搜索 · B站,分开组织

📌 订阅频道 7 个博主 · 最新上传
📺 Best Partners TV 27
🔎 搜索发现 相关度+播放量筛选 · 非订阅
🅱️ B站 AI 竞技场 按播放量

🚀 产品发布 whatships · What's Launch

📦 版本发布 tracked repos releases

🛰️ Skywork 动态

🧪 Show HN 开发者发布的新产品

💰 商业动态 · TechCrunch/VB/MIT TR

🎯 我追踪的主题(经 xlb navigator + LLM 判过价值的精选)· 图文卡片流 · 完整动态见 Reader

🎯 MCP浏览器自动化 4 精选 · 4 源更新

🆕 最新动态 4 源更新 · 同源已折叠

🎯 上下文工程精选 67 精选 · 12 源更新

🆕 最新动态 12 源更新 · 同源已折叠

🎯 嵌入式开发视频 30 精选

🎯 开源工具动态 2 精选 · 3 源更新

🆕 最新动态 3 源更新 · 同源已折叠

🎯 本地大模型 37 精选 · 15 源更新

🆕 最新动态 15 源更新 · 同源已折叠

🎯 科技分享频道 4 精选 · 30 源更新

✈️ telegram
Sliverkissの废弃文化研究所
一名实用主义技术极客的资源聚合频道。 分享AI 工具、代理软件脚本(Quantumult X / Surge / Loon / Egern) 签到与自动化、iOS 快捷指令、白嫖资讯等。 博客:blog.xn--ug8h.eu.org 线报: @sakurako0818 联系: @sliverkiss777 #AI #代理软件 #脚本分享 #快捷指令 #白嫖资讯
xlb:科技分享频道
🆕 最近动态 12
✍️ medium
🆕 最新动态 30 源更新 · 同源已折叠

🎯 自托管AI编码 15 精选 · 5 源更新

🆕 最新动态 5 源更新 · 同源已折叠