AnroMark 是什麼?
AnroMark 評估的是一台 PC 實際執行大型語言模型的能力,不是傳統 CPU 或遊戲跑分。完成本機評測後,您可將結果上傳至 Anrotec 天梯,與其他硬體配置比較表現。
天梯是什麼?
天梯是公開的硬體效能榜,依模型尺寸分成多個榜單。同一套硬體在某一尺寸榜上,以多次實測的中位吞吐(TPS)或首字延遲(TTFT)排列名次。
- 天梯名次看的是吞吐或延遲,不是本機 AI PC 分數。
- 若一次評測包含多個不同尺寸的模型,結果可分別進入多個尺寸榜。
| 榜單 | 模型規模 | 適合了解 |
|---|---|---|
| 3B 入門輕量榜 | 約 4.5B 以下 | 輕量、入門對話 |
| 7B 主流日常榜 | 約 4.5B–9B | 日常使用 |
| 14B 企業主力榜 | 約 10B–16B | 較重的日常與商務場景 |
| 20B 深度推理榜 | 約 17B–24B | 更重的模型吞吐 |
| 32B 旗艦極限榜 | 約 25B 以上 | 旗艦級極限表現 |
同一頁可切換 中位 TPS(生成速度,越高名次越好)與 首字延遲 TTFT(回應起速,越低名次越好)。
榜上每一列是一組硬體,例如 CPU 加顯示卡。天梯回答的是:「這台機器在這一類模型上,典型能跑多快?」
排名怎麼看?
請把下面兩種排名和 AI PC 分數分開看,不要當成同一個數字。
硬體天梯排名
榜單上的 #1、#2⋯
在選定的尺寸榜內,依該硬體的中位表現排序。換榜單或改看吞吐與延遲,名次都可能改變。資料較少時可能標示為初步結果。
單次評測排名
例如 12 / 340
上傳後,可查看這一次各個模型,在所有相同模型結果中的名次。例如您這次的輕量模型平均吞吐在 340 筆中排第 12。這是該次、該模型的名次。
| 您看到的 | 代表什麼 | 不代表什麼 |
|---|---|---|
| 天梯 # | 這台硬體在某尺寸榜的表現排序 | 不是 AI PC 分數 |
| 12 / 340 | 這次某個模型相對其他人的吞吐名次 | 不是天梯硬體 # |
| AI PC 分數 例如 570/1000 | 這次本機綜合評測結果,是分數不是排名 | 不直接等於天梯名次 |
本機怎麼測?
三大邏輯
真機跑模型
在您的電腦上實際載入並執行模型,量測 TPS(每秒產生多少 token,代表吞吐與生成速度)與 TTFT(第一個 token 出現要多久,代表起速與回應快慢),並參考實際 CPU 與顯示卡運作狀況,避免只有規格表、沒有真實體驗。
依模型大小分層計分
本機 AI PC 分數(滿分顯示 1000)依模型大小分為三層加權:
| 層級 | 大約規模 | 權重 |
|---|---|---|
| 輕量 | 較小模型,約 4.5B 以下 | 20% |
| 主力 | 中型模型,約 4.5B–16B | 40% |
| 重量級 | 較大模型,約 16B 以上 | 40% |
每一層綜合吞吐與起速。預設常見組合(小、中、大)可一次涵蓋三層。天梯的多個尺寸榜與本機三層分數都跟模型大小有關,但用途不同:前者比硬體榜,後者算這次綜合分。
測多少、算多少
- 只計算您有選、且測成功的模型。
- 沒測的層級不會用猜測成績填補。
- 三層都測到,才比較能代表完整的 AI PC 表現。只測一層就算該層表現很好,也不等於整機滿檔。
只測一個約 75 TPS 的輕量模型,得到約 570/1000,代表這層表現中上。合理,但不是整機完整證明,也不等於天梯一定是第幾名。上傳後可更新 3B 入門輕量榜上該硬體的中位表現與名次,並可看到該模型的單次評測名次。
如何加入天梯?
- 在本機完成 AnroMark 評測,可選 1 至 3 個模型。
- 同意上傳後,硬體與效能結果會送至 Anrotec,公開內容以硬體與效能為主。
- 每個成功的模型依大小進入對應尺寸榜。
- 可在結果頁查看單次評測名次,並在效能榜單或自選比對中比較硬體。
What is AnroMark?
AnroMark measures how well a PC actually runs large language models, not a classic CPU or gaming benchmark. After a local assessment, you can submit results to the Anrotec ladder and compare hardware.
What is the ladder?
The ladder is a public hardware leaderboard, split by model size. On each board, systems are ranked by median throughput (TPS) or time to first token (TTFT).
- Ladder position is based on speed metrics, not your AI PC Score.
- Testing several model sizes in one run can place you on more than one board.
| Board | Model size | Good for |
|---|---|---|
| 3B Entry Lightweight | about 4.5B and under | Lightweight, entry-level chat |
| 7B Everyday | about 4.5B–9B | Everyday use |
| 14B Enterprise Core | about 10B–16B | Heavier everyday and business work |
| 20B Deep Inference | about 17B–24B | Heavier model throughput |
| 32B Flagship | about 25B and above | Flagship-level limits |
Switch between median TPS, where higher is better, and TTFT, where lower is better.
Each row is a hardware setup. The ladder answers: how fast does this machine typically run models in this size class?
How to read rankings
There are two kinds of rank, plus a separate AI PC Score, which is a score rather than a place.
Ladder hardware rank
#1, #2 on a board
Ordered by that hardware's median result within one size board. Switching board or metric can change the position. Sparse data may be marked preliminary.
Single-run rank
for example 12 / 340
After upload you can see how this run's models place among all results for the same model. It belongs to that run and that model only.
| What you see | Meaning | Not the same as |
|---|---|---|
| Ladder # | Hardware rank on one size board | AI PC Score |
| 12 / 340 | This run's throughput rank for one model among others of the same model | Ladder hardware # |
| AI PC Score e.g. 570 / 1000 | This assessment's overall local score | Ladder position |
How local testing works
Three principles
Real on-device runs
Models are loaded and run on your machine to measure throughput (tokens per second, TPS) and time to first token (TTFT), alongside real CPU and GPU behaviour rather than spec sheets alone.
Size-based scoring
The AI PC Score, shown out of 1000, weights three layers by model size:
| Layer | Approximate size | Weight |
|---|---|---|
| Lighter | under about 4.5B | 20% |
| Mid | about 4.5B–16B | 40% |
| Heavier | above about 16B | 40% |
Each layer combines speed and responsiveness. A common small, medium and large set covers all three at once.
Only what you tested counts
- Only models you selected and that completed successfully are counted.
- Untested layers are not filled with guesses.
- A full three-layer run best represents overall AI PC capability.
A single lightweight model at about 75 TPS scoring near 570 / 1000 is a fair mid-range result for that layer. It is not a full-machine claim, and not an automatic ladder place. After upload it can update that size board's hardware standing and show a per-model run rank.
How to join
- Complete an AnroMark assessment locally, choosing 1 to 3 models.
- Consent to upload. Hardware and performance results are sent to Anrotec, and what is published centres on hardware and performance.
- Each successful model joins the board matching its size.
- Check your run ranks on the result page, then browse or compare hardware on the ladder.
下載 AnroMark Assessment Tool
AnroMark Assessment Tool 會在您的電腦上實際載入並執行大型語言模型,量測吞吐與首字延遲,算出本機 AI PC 分數。您同意之後,才會把結果上傳到天梯。
下載連結尚未開放。需要提前試用,請透過 Anrotec 聯絡我們。
安裝之後的流程
- 開啟工具,選擇要測試的模型,可選 1 至 3 個。建議小、中、大各一個,三層都測到。
- 開始評測。工具會在本機實際跑模型,量測 TPS 與 TTFT。
- 檢視本機結果與 AI PC 分數。
- 確認同意後上傳,即可加入對應的尺寸榜。
- 在結果頁查看單次評測名次,並回到天梯比較硬體。
- 上傳是選擇性的,沒有同意就不會送出。
- 公開內容以硬體與效能為主。
- 只有測成功的模型會計分,沒測的層級不會補分。
Download the AnroMark Assessment Tool
The AnroMark Assessment Tool loads and runs large language models on your machine, measures throughput and time to first token, and produces a local AI PC Score. Nothing is submitted to the ladder until you consent.
The download is not open yet. For early access, contact Anrotec.
What happens after install
- Open the tool and pick 1 to 3 models. One small, one medium and one large covers all three layers.
- Start the assessment. Models run locally while TPS and TTFT are measured.
- Review your local results and AI PC Score.
- Consent to upload and join the matching size boards.
- Check your run ranks on the result page, then compare hardware on the ladder.
- Uploading is optional and never happens without your consent.
- What is published centres on hardware and performance.
- Only successful runs are scored, and untested layers are not filled in.