AnroMark

加入 AnroMark 天梯 Join the AnroMark Ladder

在自己的機器上實測大型語言模型,上傳後就能與其他硬體配置同場比較。 Run large language models on your own machine, then submit the result and compare against other hardware.

AnroMark 是什麼?

AnroMark 評估的是一台 PC 實際執行大型語言模型的能力,不是傳統 CPU 或遊戲跑分。完成本機評測後,您可將結果上傳至 Anrotec 天梯,與其他硬體配置比較表現。

天梯是什麼?

天梯是公開的硬體效能榜,依模型尺寸分成多個榜單。同一套硬體在某一尺寸榜上,以多次實測的中位吞吐(TPS)首字延遲(TTFT)排列名次。

請分開理解
  • 天梯名次看的是吞吐或延遲,不是本機 AI PC 分數
  • 若一次評測包含多個不同尺寸的模型,結果可分別進入多個尺寸榜
榜單模型規模適合了解
3B 入門輕量榜約 4.5B 以下輕量、入門對話
7B 主流日常榜約 4.5B–9B日常使用
14B 企業主力榜約 10B–16B較重的日常與商務場景
20B 深度推理榜約 17B–24B更重的模型吞吐
32B 旗艦極限榜約 25B 以上旗艦級極限表現

同一頁可切換 中位 TPS(生成速度,越高名次越好)與 首字延遲 TTFT(回應起速,越低名次越好)。

榜上每一列是一組硬體,例如 CPU 加顯示卡。天梯回答的是:「這台機器在這一類模型上,典型能跑多快?」

排名怎麼看?

請把下面兩種排名AI PC 分數分開看,不要當成同一個數字。

硬體天梯排名

榜單上的 #1、#2⋯

在選定的尺寸榜內,依該硬體的中位表現排序。換榜單或改看吞吐與延遲,名次都可能改變。資料較少時可能標示為初步結果。

單次評測排名

例如 12 / 340

上傳後,可查看這一次各個模型,在所有相同模型結果中的名次。例如您這次的輕量模型平均吞吐在 340 筆中排第 12。這是該次、該模型的名次。

您看到的代表什麼不代表什麼
天梯 #這台硬體在某尺寸榜的表現排序不是 AI PC 分數
12 / 340這次某個模型相對其他人的吞吐名次不是天梯硬體 #
AI PC 分數
例如 570/1000
這次本機綜合評測結果,是分數不是排名不直接等於天梯名次

本機怎麼測?

三大邏輯

1

真機跑模型

在您的電腦上實際載入並執行模型,量測 TPS(每秒產生多少 token,代表吞吐與生成速度)與 TTFT(第一個 token 出現要多久,代表起速與回應快慢),並參考實際 CPU 與顯示卡運作狀況,避免只有規格表、沒有真實體驗。

2

依模型大小分層計分

本機 AI PC 分數(滿分顯示 1000)依模型大小分為三層加權:

層級大約規模權重
輕量較小模型,約 4.5B 以下20%
主力中型模型,約 4.5B–16B40%
重量級較大模型,約 16B 以上40%

每一層綜合吞吐起速。預設常見組合(小、中、大)可一次涵蓋三層。天梯的多個尺寸榜與本機三層分數都跟模型大小有關,但用途不同:前者比硬體榜,後者算這次綜合分。

3

測多少、算多少

  • 只計算您有選、且測成功的模型。
  • 沒測的層級不會用猜測成績填補。
  • 三層都測到,才比較能代表完整的 AI PC 表現。只測一層就算該層表現很好,也不等於整機滿檔。
例子

只測一個約 75 TPS 的輕量模型,得到約 570/1000,代表這層表現中上。合理,但不是整機完整證明,也不等於天梯一定是第幾名。上傳後可更新 3B 入門輕量榜上該硬體的中位表現與名次,並可看到該模型的單次評測名次。

如何加入天梯?

  1. 在本機完成 AnroMark 評測,可選 1 至 3 個模型。
  2. 同意上傳後,硬體與效能結果會送至 Anrotec,公開內容以硬體與效能為主。
  3. 每個成功的模型依大小進入對應尺寸榜。
  4. 可在結果頁查看單次評測名次,並在效能榜單或自選比對中比較硬體。
先看看天梯榜單

What is AnroMark?

AnroMark measures how well a PC actually runs large language models, not a classic CPU or gaming benchmark. After a local assessment, you can submit results to the Anrotec ladder and compare hardware.

What is the ladder?

The ladder is a public hardware leaderboard, split by model size. On each board, systems are ranked by median throughput (TPS) or time to first token (TTFT).

Keep these apart
  • Ladder position is based on speed metrics, not your AI PC Score.
  • Testing several model sizes in one run can place you on more than one board.
BoardModel sizeGood for
3B Entry Lightweightabout 4.5B and underLightweight, entry-level chat
7B Everydayabout 4.5B–9BEveryday use
14B Enterprise Coreabout 10B–16BHeavier everyday and business work
20B Deep Inferenceabout 17B–24BHeavier model throughput
32B Flagshipabout 25B and aboveFlagship-level limits

Switch between median TPS, where higher is better, and TTFT, where lower is better.

Each row is a hardware setup. The ladder answers: how fast does this machine typically run models in this size class?

How to read rankings

There are two kinds of rank, plus a separate AI PC Score, which is a score rather than a place.

Ladder hardware rank

#1, #2 on a board

Ordered by that hardware's median result within one size board. Switching board or metric can change the position. Sparse data may be marked preliminary.

Single-run rank

for example 12 / 340

After upload you can see how this run's models place among all results for the same model. It belongs to that run and that model only.

What you seeMeaningNot the same as
Ladder #Hardware rank on one size boardAI PC Score
12 / 340This run's throughput rank for one model among others of the same modelLadder hardware #
AI PC Score
e.g. 570 / 1000
This assessment's overall local scoreLadder position

How local testing works

Three principles

1

Real on-device runs

Models are loaded and run on your machine to measure throughput (tokens per second, TPS) and time to first token (TTFT), alongside real CPU and GPU behaviour rather than spec sheets alone.

2

Size-based scoring

The AI PC Score, shown out of 1000, weights three layers by model size:

LayerApproximate sizeWeight
Lighterunder about 4.5B20%
Midabout 4.5B–16B40%
Heavierabove about 16B40%

Each layer combines speed and responsiveness. A common small, medium and large set covers all three at once.

3

Only what you tested counts

  • Only models you selected and that completed successfully are counted.
  • Untested layers are not filled with guesses.
  • A full three-layer run best represents overall AI PC capability.
Example

A single lightweight model at about 75 TPS scoring near 570 / 1000 is a fair mid-range result for that layer. It is not a full-machine claim, and not an automatic ladder place. After upload it can update that size board's hardware standing and show a per-model run rank.

How to join

  1. Complete an AnroMark assessment locally, choosing 1 to 3 models.
  2. Consent to upload. Hardware and performance results are sent to Anrotec, and what is published centres on hardware and performance.
  3. Each successful model joins the board matching its size.
  4. Check your run ranks on the result page, then browse or compare hardware on the ladder.
Browse the ladder first
AnroMark · 由 Anrotec 維運。天梯名次依中位吞吐或首字延遲排序,與本機 AI PC 分數是不同的數字。 Operated by Anrotec. Ladder position is ranked by median throughput or TTFT, which is a different number from the local AI PC Score.
anrotec.ai · anromark.ai · 效能榜單Benchmarks · © Anrotec Inc.