⚡ GitHub Pulse

生成于 2026-09-21 18:25 UTC · 追踪 3,000 仓库 · 8 多源共振
🔍 搜索中 · 显示所有标签页的匹配项 · 按 Esc 清除
📊 深度分析(2026-09-20 20:40 北京)基于更早的数据快照,数据已刷新,已自动隐藏过时结论 —— 待重新生成分析后显示。

💎 高价值精选 AI 判断 · 跨源挖掘

📰 最新快讯

⭐ 多源共振

jaredpalmer/kev ×5 🆕 new
githubxsocialboardhn
+1,084★/d 活跃开发 official #19
tiny Jev-like model built on top of Qwen2.5-0.5B you can train and run on your MacBook
🔺 @jaredpalmer 首发 · 47h 前
githubxsocialboard
+5,034★/d 活跃开发 official #1
🔺 @clxymox 首发 · 37h 前
githubxsocialboard
+3,309★/d 早期·低活动 official #2
🔺 @betterhn20 首发 · 16h 前
githubxsocialboard
+2,435★/d 早期·低活动 official #3
354 条循证建议,覆盖长寿防病、急救、省钱理财、法律红线、失业与工伤、医保社保、恋爱婚育、出国与技能。每条写明成本、收益、证据等级和原始出处,只引期刊论文与官方文件。
🔺 @ForestGrahxu 首发 · 66h 前
google/ax ×4 ↺ 1d
githubxsocialboard
+2,421★/d 早期·低活动 official #6
Google's open agentic orchestrator
githubxsocialboard
+1,664★/d 早期·低活动 official #4
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
🔺 @mizorewww 首发 · 23h 前
stablyai/orca ×4 ↺ 2d
githubxsocialboard
+946★/d 活跃开发 official #15
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
🔺 @tianma_if 首发 · 61h 前
githubxsocialboard
+820★/d 早期·低活动 official #16
A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
🔺 @MaciejLukianski 首发 · 41h 前

🔥 动量榜

#Repo7d+1d★7d★质地官方X
1 NandhaKishorM/laya ↺ 2d
Python
🔺 @clxymox 首发 · 37h 前
+5,034 8,998 活跃开发 #1 19×
2 zai-org/ZCode 🆕 new
TypeScript
+4,473 4,473 早期·低活动 41×
3 browser-use/jev-ultrafast ↺ 2d
Python
🔺 @betterhn20 首发 · 16h 前
+3,309 14,916 早期·低活动 #2 23×
4 eternity4719/HowToLiveBetter ↺ 2d
HTML · 354 条循证建议,覆盖长寿防病、急救、省钱理财、法律红线、失业与工伤、医保社保、恋爱婚育、出国与技能。每条写明成本、收
🔺 @ForestGrahxu 首发 · 66h 前
+2,435 10,753 早期·低活动 #3 24×
5 google/ax ↺ 1d
Go · Google's open agentic orchestrator
+2,421 3,306 早期·低活动 #6 12×
6 Albert-Weasker/niubigeo 🆕 new
TypeScript
+2,135 2,252 早期·低活动
7 mizorewww/laya-mlx ↺ 1d
Python · Native MLX runtime for Laya typed decision models — 7–14 ms
🔺 @mizorewww 首发 · 23h 前
+1,664 3,294 早期·低活动 #4
8 jaredpalmer/kev 🆕 new
Python · tiny Jev-like model built on top of Qwen2.5-0.5B you can tra
🔺 @jaredpalmer 首发 · 47h 前
+1,084 1,928 活跃开发 #19 11×
9 deepseek-ai/deepseek-harness ↺ 2d
TypeScript
🔺 @the_osps 首发 · 31h 前
+1,022 8,243 早期·低活动 11×
10 stablyai/orca ↺ 2d
TypeScript · Orca is the ADE for working with a fleet of parallel agents.
🔺 @tianma_if 首发 · 61h 前
+946 5,792 活跃开发 #15 22×
11 affaan-m/ECC ↺ 2d
JavaScript
🔺 @BlockInsight214 首发 · 10h 前
+857 6,179 活跃开发 22×
12 cloudflare/security-audit-skill ↺ 2d
JavaScript · A coding-agent skill for multi-phase security audits with in
🔺 @MaciejLukianski 首发 · 41h 前
+820 14,896 早期·低活动 #16 34×
13 pacifio/atlas 🆕 new
Rust
+817 1,231 活跃开发 10×
14 bilawalsidhu/gods-eye-view ↺ 2d
🔺 @LFrefman 首发 · 57h 前
+770 6,836 活跃开发 24×
15 dexmal/dexbotic ↺ 1d
Python
+752 1,655 早期·低活动
16 Open-Dev-Society/OpenStock ↺ 2d
TypeScript · OpenStock is an open-source alternative to expensive market
+727 3,255 早期·低活动 #8 15×
17 tt-a1i/archify ↺ 2d
JavaScript
🔺 @xiaomovps 首发 · 59h 前
+721 6,768 活跃开发 13×
18 Mak5er/AirCard 🆕 new
Swift
🔺 @PAPERonNet 首发 · 66h 前
+707 1,100 活跃开发
19 alibaba/open-code-review ↺ 2d
Go
🔺 @shao__meng 首发 · 71h 前
+707 13,467 活跃开发 23×
20 hypit-ai/hypit ↺ 2d
TypeScript
🔺 @cccyd_qwq 首发 · 31h 前
+685 11,927 活跃开发 32×

🗞️ Hacker News

🛠️ 技术源 · GitHub Trending/Lobsters

🐧 LINUX DO

💬 V2EX

👽 Reddit

📥 AI 博客 · Newsletter

📢 电报精选

📚 科技周刊 新项目/工具自荐

🎯 Alpha 账号 X 上最早带火仓库的人

作者leadslead率仓库数
@shanyanggm220.7123
@xzbx888130.878
@shaw_stone7383290.697
@GitTrend0x80.738
@LFrefman80.479
@the_osps70.886
@xfubot70.548
@the_rza_70.448
@Sn0wbrave60.556
@clxymox60.48
@shao__meng60.66
@vintcessun60.3513
@iasg100460.2716
@FrontieraTechIT50.637
@bilawalsidhu511
@DataChaz50.835
@RepoGems50.834
@neil_xbt50.633
@key_indie50.832
@LoveAIbrain514
@jasontopia40.83
@QingQ7740.317
@seekjourney40.673
@ctatedev40.82
@ClaudeCodeLog411

🏆 各领域最强模型

文本 / 对话
Claude Fable 5.1
Anthropic · 53.4
图像生成
GPT Image 2.5 Flare
OpenAI · 1188
图像编辑
GPT Image 2.5 Sunburst
OpenAI · 1176
文生视频
Wan 3.0
Alibaba · 1336
图生视频
Gemini Omni Flash
Google · 1369
语音合成
Sonic 3.6
Cartesia · 1276

🏆 能力排行榜 Artificial Analysis

文本 / 对话 Intelligence Index
  1. 1Claude Fable 5.153.4
  2. 2GPT-6 Astra52.7
  3. 3Claude Opus 550.8
  4. 4Claude Fable 549.6
  5. 5Muse Spark 1.348.1
  6. 6GPT-5.6 Sol47
  7. 7Grok 4.7 🆕46.4
  8. 8Qwen3.8 Max45.4
  9. 9GLM-5.344.8
  10. 10Grok 4.644.3
  11. 11Step 5 Preview43.7
  12. 12Kimi K343.6
图像生成 Text→Image Arena Elo
  1. 1GPT Image 2.5 Flare1188
  2. 2GPT Image 2.5 Sunburst1182
  3. 3GPT Image 21171
  4. 4Grok Imagine Image 2.01154
  5. 5MAI-Image-2.61147
  6. 6Reve 2.11129
  7. 7Nano Banana 21122
  8. 8Muse Image1111
  9. 9GPT Image 1.51102
  10. 10MAI-Image-2.51102
  11. 11Nano Banana Pro1100
  12. 12MAI-Image-2.6-Flash1099
图像编辑 Image-Editing Arena Elo
  1. 1GPT Image 2.5 Sunburst1176
  2. 2GPT Image 2.5 Flare1155
  3. 3MAI-Image-2.61132
  4. 4MAI-Image-2.6-Flash1122
  5. 5GPT Image 21121
  6. 6Muse Image1115
  7. 7MAI-Image-2.51113
  8. 8MAI-Image-2.5-Pro1106
  9. 9Seedream 5.0 Pro1106
  10. 10Nano Banana 21105
  11. 11Grok Imagine Image 2.01104
  12. 12GPT Image 1.51104
文生视频 Text→Video Arena Elo
  1. 1Wan 3.01336
  2. 2Gemini Omni Flash1330
  3. 3MiniMax H31302
  4. 4HappyHorse-1.01287
  5. 5HappyHorse-1.11272
  6. 6Dreamina Seedance 2.0 720p1259
  7. 7Wan2.7-2606121243
  8. 8grok-imagine-video1235
  9. 9Kling 3.0 Omni 1080p1230
  10. 10PixVerse V5.61230
  11. 11PixVerse V61230
  12. 12Kling 3.0 1080p1230
图生视频 Image→Video Arena Elo
  1. 1Gemini Omni Flash1369
  2. 2Wan 3.01361
  3. 3Bach 1.0 Pro1359
  4. 4MiniMax H31354
  5. 5PixVerse V61337
  6. 6Dreamina Seedance 2.0 720p1336
  7. 7grok-imagine-video-1.51329
  8. 8grok-imagine-video1326
  9. 9HappyHorse-1.11312
  10. 10Kling 2.5 Turbo 1080p1296
  11. 11HappyHorse-1.01293
  12. 12Vidu Q3 Pro1290
语音合成 Text→Speech Arena Elo
  1. 1Sonic 3.61276
  2. 2Qwen-Audio-3.0-TTS-Plus1260
  3. 3Realtime TTS-21247
  4. 4Simba 3.21240
  5. 5Luna TTS1231
  6. 6Realtime TTS-2 Flash1215
  7. 7StepAudio 2.5 TTS1209
  8. 8Breeze TTS 21205
  9. 9Gemini 3.1 Flash TTS1201
  10. 10v3 Conversational1197
  11. 11Sonic 3.51184
  12. 12Lightning V3.1 Pro1179

🥇 综合能力榜 Benchmark Heaven · 7 榜合一(AA+Epoch ECI+DesignArena)

#模型综合分$/1MBenchmaxxing
1Claude Fable 5.1 Anthropic 🏅前沿98.4$13.640.62
2GPT 6 Astra OpenAI97.7$13.64-6.41
3Claude Opus 5 Anthropic96.3$6.82-5.96
4Claude Fable 5 Anthropic95.6$13.64-4.46
5Muse Spark 1.3 Meta 🏅前沿89.3$1.525.8
6GPT 5.6 Sol OpenAI 🏅前沿95.2$2.73-2.64
7Qwen3.8 Max 0902 Alibaba87.7$2.36
8GLM 5.3 Z.ai 开源 🏅前沿85.9$0.921.63
9Grok 4.6 xAI84.6$2.365.33
10Step 5 Preview StepFun 🏅前沿87.4$1.15
11Kimi K3 Moonshot AI 开源 🏅前沿89.5$2.323.73
12GPT 5.6 Terra OpenAI82.2$2.91-3.75
13GLM 5.3 Flash Z.ai 开源 🏅前沿77.5$0.09-1.15
14Claude Opus 4.8 Anthropic88.4$6.82-2.34
15Gemini 3.8 Flash Google81.4$1.026.3

Benchmaxxing:BH 的指标,+ 值越高表示该模型在"公开基准"上比"不可训练的封闭题"排名更靠前(BH 定义与计算,非本站判断)。

💰 性价比 / 最省钱 跨 provider 最低价 · 10:1 blended · 综合分≥60

模型$/1M综合分最便宜 provider
DeepSeek V4 Flash 0731 🏅前沿 🔒不训练 开源$0.0574.4Relace
Agnes 3.0 Flash$0.0671.2
GLM 5.3 Flash 🏅前沿 开源$0.0977.5GMICloud
Agnes 2.5 Pro Beta$0.1271.2
DeepSeek V4.1 Flash 🔒不训练 开源$0.1569.9Morph
Qwen3.8 Flash Next 🏅前沿 开源$0.1882Alibaba
Qwen3.8 27B 开源$0.2569.8Darkbloom
GPT 5.6 Luna 🔒不训练$0.2979.5Azure AI Foundry
Solar Pro 4$0.3863.1
Agnes 2.5 Pro Alpha 开源$0.4962.2
DeepSeek V4 Flash Vision$0.5266.4
Inkling Small 🔒不训练 开源$0.5264DeepInfra
Apodex 1.1$0.5571.2
GLM 5.2 🔒不训练 开源$0.6876.1DeepInfra
Quasar 438B$0.7166.1

跨 95 家 provider(含 21 家 🇪🇺 EU、53 家 🔒不训练)· 数据 2026-09-21

🏅 性价比前沿 没有更便宜的模型能在能力上胜过它们(Pareto)

🧠 最新发布

2026-09-18
GLM 5.3 FlashX
Zhipu AI
2026-09-18
Ternary Bonsai 2 27B
PrismML
2026-09-17
KAT-Coder-Pro V2.5
Kwaipilot
2026-09-17
Gemini Omni Flash Preview
Google
2026-09-17
Venice Uncensored
Venice
2026-09-17
Nano Banana 2 Lite
Google
2026-09-17
Pareto
Unbiased
2026-09-17
Qwen3.8 Omni Flash
Alibaba Cloud

🧠 模型发布时间线

2026-09-18
GLM 5.3 FlashX
Zhipu AI
aimlapi
2026-09-18
Ternary Bonsai 2 27B
PrismML
aimlapi
2026-09-17
KAT-Coder-Pro V2.5
Kwaipilot
aimlapi
2026-09-17
Gemini Omni Flash Preview
Google
aimlapi
2026-09-17
Venice Uncensored
Venice
aimlapi
2026-09-17
Nano Banana 2 Lite
Google
aimlapi
2026-09-17
Pareto
Unbiased
aimlapi
2026-09-17
Qwen3.8 Omni Flash
Alibaba Cloud
aimlapi
2026-09-16
Union Alpha
Stealth
aimlapi
2026-09-15
Jev 1.13
TypeSafe AI
aimlapillmgateway
2026-09-14 · ★
Grok 4.7
X AI
aimlapiopperllmgateway
2026-09-12
Schematron V2 Turbo
Inference.net
aimlapi
2026-09-12
Schematron V2 Small
Inference.net
aimlapi
2026-09-11
Fugu Ultra v2.0
Sakana AI
llmgateway
2026-09-11
Kimi K2.8 Preview
Moonshot AI
llmstats
2026-09-11
Atria Dawn Preview
Shanghai AI Laboratory
llmstatsllmgateway
2026-09-11
Fugu Ultra v2
Sakana AI
aimlapi
2026-09-11
Fugu Max
Sakana AI
aimlapillmgateway
2026-09-10 · ★
Ling 3.0 Flash VL
inclusionAI
aimlapiopper
2026-09-10 · ★
DeepSeek V4.1 Flash
DeepSeek AI
aimlapillmstatsopperllmgateway
2026-09-10
DeepSeek Chat (V4.1 Flash)
DeepSeek AI
aimlapi
2026-09-08
GPT Image 2.5 Sunburst
Open AI
aimlapillmgateway
2026-09-08
GPT Image 2.5 Flare
Open AI
aimlapillmgateway
2026-09-08
Mercury 2.5
Inception
aimlapi

📄 论文 PwC + arXiv

🤗 HF 采用榜 下载/点赞

Modeldownloadslikespipeline
convaiinnovations/laya01,613text-classification
prism-ml/Ternary-Bonsai-2-27B-gguf2,227,8791,697text-generation
Qwen/Qwen-Image-2.16,5231,336text-to-image
XingChen-AGI/Xing4.0-29B-A4B18,3941,071text-generation
deepseek-ai/DeepSeek-V4.1-Flash512,1203,506image-text-to-text
Qwen/Qwen3.8-27B7,153,23815,952image-text-to-text
harshatheg/Qwen-2.5-1B-RLCD0505text-generation
ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF1,292,4711,525image-text-to-text
abenzerps/Qwen-Image-2.1-GGUF33,232492text-to-image
m-a-p/YuE2-3B18,759938text-to-audio
Lightricks/LTX-2.51,626,7424,639image-to-video
Comfy-Org/Qwen-Image-2.1535,365417
ukisai/Swift-Qwen3.8-27b16,514525image-text-to-text
AlexWortega/openjev0399text-classification
unsloth/Qwen3.8-27B-GGUF7,039,0064,464
Altworld/Hemmingway-1834323text-generation
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF1,348,7121,036image-text-to-text
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit36,744307text-generation
Qwen/Qwen3.8-Flash-Next774,7785,538image-text-to-text
ukisai/Swift-Qwen3.8-27B-GGUF144,372341image-text-to-text
openbmb/MiniCPM5-2B460,5331,640text-generation
MiniMaxAI/MiniMax-H34,046,9175,564image-text-to-video
TaichuAI/ZDTaichu5.0-9B5,078218image-text-to-text
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF53,094205image-text-to-text
netease-youdao/Confucius4-R2T21,864216automatic-speech-recognition
TokenRhythm/NeoHorse-1-9B12,260988text-generation
WarmBloodAban/Minimax-h3_Singularity268,296592image-to-video
internlm/Atria-Dawn-Preview978223
zai-org/GLM-5.3-Flash3,309,5652,513image-text-to-text
Cactus-Compute/needle346,399161text-generation

⚡ System One 决策模型 420 项目 · Jev/TypeSafe 生态 · 快决策(非推理)

SDK & Decision Frameworks 75Security & Guardrails 37High-Frequency & Simulation 35Browser & OS Action 34Routing & Cost Optimization 34Evaluation & Observability 29CLI & Pipelines 26Domain & Vertical Tools 26Data & Search 26Context GC & Filter 25MCP & Integrations 22Creative Tools 15Codebase & Graph Pathfinding 12Decision Tools 12SDK & Integrations 6Voice & Conversation 4Classification & Taxonomy 2
项目类别Jev 决策点
langchain langchain-ai146,785SDK & IntegrationsSubmits binary, categorical and ordered-score questions and returns typed answers with probabilities.
ai-hedge-fund virattt63,644Domain & Vertical ToolsConverts strategy questions to System One requests and normalizes native answers to the project’s result format.
litellm BerriAI59,309Routing & Cost OptimizationMaps requests to configured complexity classes that drive backend routing.
oh-my-pi can135732,249Routing & Cost OptimizationSends agent state and typed questions to Jev and parses structured answers.
jev-model-router davila730,871Routing & Cost OptimizationEvaluates task tier, reasoning needs and production risk; local policy maps results to invocation settings.
composio ComposioHQ30,271SDK & Decision FrameworksTurns tool or action conditions into structured questions and passes Jev answers to local invocation logic.
ai vercel26,869SDK & Decision FrameworksMaps choice, score, and yes/no questions to TypeSafe System One requests and parses typed results.
cua trycua25,476Browser & OS ActionReads DOM or supported visual-region descriptions and returns a supplied candidate action ID.
pydantic-ai pydantic20,088SDK & IntegrationsConverts supported structured output fields into typed Jev questions and maps answers back to the output model.
eliza elizaOS19,404SDK & Decision FrameworksOnly an explicit systemOne call sends state and questions, returning validated typed answers.
langchainjs langchain-ai18,215SDK & Decision FrameworksUses invoke to call TypeSafe and parse choice, noul, score and probability fields.
json-render vercel-labs17,773Creative ToolsEvaluates component configurations through Vercel AI Gateway, then composeSpec assembles the UI specification.
jev-ultrafast browser-use14,237Browser & OS ActionChooses an action and its matching DOM element in one request; a text model generates input text.
openchamber openchamber10,219Routing & Cost OptimizationJev selects a task category; local category mappings determine the model configuration.
rig-typesafeai 0xPlaygrounds8,691SDK & Decision FrameworksSends application state and questions to Jev and parses Choice, Score or Noul answers.
firstmate kunchenguid6,894Routing & Cost OptimizationSends the task brief and candidate rules to Jev, then resolves execution profiles with confidence and local conditions.
fast-jev-compaction tamaratran5,789Context GC & FilterSeparately judges whether a tool call and its full output are still needed; code keeps, truncates or drops them.
agentgateway agentgateway4,954Security & GuardrailsJev scores jailbreaks, harmful content and secret disclosure; thresholds or evaluation errors reject requests.
latitude-llm latitude-dev4,664Evaluation & ObservabilityJudges which checks apply and can add checks when thresholds and rate limits permit.
ax ax-llm2,935SDK & IntegrationsMaps supported signatures to Jev questions or sends native System One requests.
kev jaredpalmer1,734High-Frequency & SimulationAttaches a parallel decision head to an open 0.5B model to answer discrete questions directly from token activations.
NanoJev TianyuCodings1,704High-Frequency & SimulationEvaluates multiple questions and dynamic candidate spaces concurrently in a single forward pass, logging navigation choices.
jev-trader jarrodwatts1,699Domain & Vertical ToolsIn Jev mode, order-book judgments feed code that simulates fills or submits configured post-only limit orders.
jev-desktop lahfir1,387Browser & OS ActionJev selects a target and action and estimates presence and risk; local policy decides whether to execute.
vellum-assistant vellum-ai1,293MCP & IntegrationsSubmits state and question bundles to System One and returns structured answers to the Assistant.
celesto CelestoAI951Codebase & Graph PathfindingJudges whether a finding was introduced by the change, is supported and merits a fix.
atomic bastani-inc809Routing & Cost OptimizationSends predefined questions to Jev and decodes answers for callers; regular models still generate code.
typesafe-computer-use awlevin706Browser & OS ActionSelects the next step from deterministically extracted controls and actions before desktop execution.
aiavatarkit uezo678Voice & ConversationAssesses utterance completeness and whether the user is likely to continue speaking.
kody kentcdodds660Data & SearchSends a Score question per candidate, reorders and drops low scores; the model id is typesafe/jev.
Agent AgentiLoop619Security & GuardrailsAdds a destructive-risk judgment after local shell checks and refuses commands above the configured threshold.
req_llm agentjido580SDK & Decision FrameworksSends state and questions, normalizes answers and retains the raw provider response.
omg.dev BennyKok533Browser & OS ActionChooses controls and checks completion or blockage before the test runner operates the UI.
Jev-cu Sac-Y522Browser & OS ActionChooses targets and actions and assesses completion and risk; local policy controls execution or confirmation.
foreman thruwire451CLI & PipelinesAsyncTypeSafeClient.system_one with default jev-latest sends Noul questions for supervision.
Jev Review devagrawal09446Codebase & Graph PathfindingJudges risk, files, evidence regions, mechanisms and severity before rule-based reviewer routing.
simple-jev featherless-ai432SDK & Decision FrameworksExtracts log-probabilities of candidate tokens from model vocabulary logits, formatting them into standard Jev responses.
vexjoy-agent notque421Routing & Cost OptimizationAfter deterministic routing guards, Jev judges the remaining candidates and required workflow components.
jev-search superagents-lab355Data & SearchJudges search intent, source and time settings, and relevance of individual results.
jev-experiments dabit3353Security & GuardrailsJudges risks such as exposed credentials or destructive changes; local rules warn or block a commit.
🏅 JevBench 决策榜 JevBench v1 · 242 决策/模型 · 私有留出集 · 准确率 + 校准(Brier↓ 越低越准)
#系统准确率Brier↓类型
1Jev 1.13.0 (TypeSafe AI) TypeSafe AI 开源100%0.0028jev
2openjev-sglang (Qwen3.6-35B-A3B on SGLang) ekzhang 开源100%0.0305jev-rebuild
3system-one-open (Gemma 4 E2B LoRA on an L4) mithalouni 开源96%0.0851jev-rebuild
4open-alternative-jev (Qwen3.5-4B, HF Space) IkerMoel 开源jev-rebuild
5open-jev-deberta-v3-large (local CPU) Kotoba Labs 开源67%0.4769jev-rebuild
6GPT-5.6 Luna (low reasoning effort) OpenAI 开源100%0.0003llm-baseline
7Gemini 3.1 Flash-Lite Google 开源100%0.0017llm-baseline
8DeepSeek V4.1 Flash (thinking default) DeepSeek 开源100%0.0001llm-baseline
9Qwen3.8 27B (Chutes TEE) Qwen / Chutes 开源100%0.0001llm-baseline

JevBench(BH):把"决策模型"(Jev/System One)与通用 LLM 放在同一批决策任务上比准确率与概率校准。

🅱️ B站 AI 无限竞技场 18 模型 · 夺冠率

#模型夺冠率冠/测
1GPT-6 Astra OpenAI55%12/22
2Claude Fable 5.1 Anthropic50%6/12
3GLM-5.3 Z.ai16%3/19
3GPT-5.6 Sol OpenAI16%3/19
5Claude Fable 5 Anthropic27%3/11
6Kimi K3 Moonshot10%2/21
7Claude Opus 5 Anthropic11%2/19
8DeepSeek-V4-Flash DeepSeek13%2/16
9DeepSeek-V4-Pro DeepSeek5%1/21
10Qwen3.8-Max Alibaba6%1/18
11DeepSeek V4.1 Flash DeepSeek8%1/13
12Gemini 3.8 Flash Google8%1/12
13Hy 4 Tencent14%1/7
14Gemini 3.7 Flash Google20%1/5
15GPT-5.6 Terra OpenAI25%1/4
15Seed-2.0 pro ByteDance25%1/4
17Doubao-Seed-Evolving ByteDance33%1/3
18Seed-2.0 Mini ByteDance100%1/1
👤 AI UP主 从赛题发现,点击直达主页
🎯 赛题

🎬 AI 视频 订阅 · 搜索 · B站,分开组织

📌 订阅频道 5 个博主 · 最新上传
📺 Best Partners TV 15
📺 AI超元域 12
🔎 搜索发现 相关度+播放量筛选 · 非订阅
🅱️ B站 AI 竞技场 按播放量

🚀 产品发布 whatships · What's Launch

📦 版本发布 tracked repos releases

🛰️ Skywork 动态

🧪 Show HN 开发者发布的新产品

💰 商业动态 · TechCrunch/VB/MIT TR

🎯 我追踪的主题(经 xlb navigator + LLM 判过价值的精选)· 图文卡片流 · 完整动态见 Reader

🎯 MCP浏览器自动化 4 精选 · 4 源更新

🆕 最新动态 4 源更新 · 同源已折叠

🎯 上下文工程精选 27 精选 · 10 源更新

🆕 最新动态 10 源更新 · 同源已折叠

🎯 嵌入式开发视频 30 精选

🎯 开源工具动态 2 精选 · 3 源更新

🆕 最新动态 3 源更新 · 同源已折叠

🎯 本地大模型 35 精选 · 13 源更新

🆕 最新动态 13 源更新 · 同源已折叠

🎯 科技分享频道 4 精选 · 30 源更新

🆕 最新动态 30 源更新 · 同源已折叠

🎯 自托管AI编码 12 精选 · 3 源更新

🆕 最新动态 3 源更新 · 同源已折叠