<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet href="/rss.xsl" type="text/xsl"?>
<rss version="2.0">
  <channel>
    <title>IT社区推荐资讯 - ITIndex.net</title>
    <link>https://itindex.net/</link>
    <description>IT社区推荐资讯 - ITIndex.net</description>
    <language>zh</language>
    <copyright>https://itindex.net/</copyright>
    <generator>https://itindex.net/</generator>
    <docs>http://backend.userland.com/rss</docs>
    <image>
      <url>https://itindex.net/images/logo.gif</url>
      <title>IT社区推荐资讯 - ITIndex.net</title>
      <link>https://itindex.net/</link>
    </image>
    <item>
      <title>小米AI Cube 原型机-本地AI模型部署选择</title>
      <link>https://itindex.net/detail/63291-%E5%B0%8F%E7%B1%B3-ai-cube</link>
      <description>&lt;div&gt;    &lt;div&gt;      &lt;div&gt;        &lt;div&gt;          &lt;h2&gt;小米刚刚展示了其 AI Cube 原型机，这可能成为来自中国的严肃 GB10 竞争对手 👀

- 3 款定制芯片：Xring O3、O100、D100&lt;/h2&gt;     &lt;h2&gt;
- 200 TOPS NPU&lt;/h2&gt;     &lt;h2&gt;
- 1.22 TB/s AI 内存带宽&lt;/h2&gt;     &lt;h2&gt;
- 最高 160GB 统一内存&lt;/h2&gt;     &lt;h2&gt;
- 150W 持续功率&lt;/h2&gt;     &lt;h2&gt;
- 120B 模型本地运行&lt;/h2&gt;     &lt;h2&gt;

Xring O100：1.22TB/s + 330 t/s 在 150w AI 盒子上真是 🔥

一旦它上市，绝对会像热饼一样热卖。&lt;/h2&gt;     &lt;h2&gt;Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀

- 3 custom chips: Xring O3, O100, D100
- 200 TOPS NPU
- 1.22 TB/s AI memory bandwidth
- Up to 160GB unified memory
- 150W sustained power
- 120B models running locally

Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥

Once it hit&amp;apos;s the marked, going to sell like hot cakes.&lt;/h2&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;  &lt;div&gt;         &lt;div&gt;           &lt;div&gt;             &lt;div&gt;               &lt;div&gt;                 &lt;div&gt;                   &lt;div&gt;                     &lt;div&gt;                       &lt;div&gt;                         &lt;div&gt;                           &lt;div&gt;                             &lt;div&gt;                               &lt;div&gt;                                 &lt;img alt="" height="608" src="https://pbs.twimg.com/media/HQepUd0bkAAI5L_?format=webp&amp;name=large" width="1080"&gt;&lt;/img&gt;&lt;/div&gt;                               &lt;a href="https://x.com/ItsmeAjayKV/status/2091827310160908400/photo/1"&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;                         &lt;div&gt;                           &lt;div&gt;                             &lt;div&gt;                               &lt;div&gt;                                 &lt;img alt="" height="1145" src="https://pbs.twimg.com/media/HQepzk2bsAAODlD?format=webp&amp;name=large" width="900"&gt;&lt;/img&gt;&lt;/div&gt;                               &lt;a href="https://x.com/ItsmeAjayKV/status/2091827310160908400/photo/2"&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;         &lt;div&gt;    &lt;br /&gt;            &lt;div&gt;     &lt;br /&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63291-%E5%B0%8F%E7%B1%B3-ai-cube</guid>
      <pubDate>Wed, 07 Oct 2026 08:13:57 CST</pubDate>
    </item>
    <item>
      <title>文本分类语言模型-Language Models for Text Classification: From Bag-of-Words to Jev</title>
      <link>https://itindex.net/detail/63290-%E6%96%87%E6%9C%AC-%E5%88%86%E7%B1%BB-%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B</link>
      <description>&lt;div&gt;    &lt;div&gt;      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;            &lt;div&gt;              &lt;h1&gt;        &lt;br /&gt;&lt;/h1&gt;              &lt;h3&gt;词袋模型、循环神经网络、卷积神经网络、Transformer 模型、Jev 类 API 和校准的可视化指南&lt;/h3&gt;       &lt;h3&gt;A Visual Guide to Bag-of-Words, RNNs, CNNs, Transformers, Jev-like APIs, and Calibration&lt;/h3&gt;              &lt;div&gt;                &lt;div&gt;                  &lt;div&gt;                    &lt;div&gt;                      &lt;div&gt;                        &lt;div&gt;                          &lt;a href="https://substack.com/@rasbt"&gt;                            &lt;div&gt;                              &lt;div&gt;                                &lt;img alt="Sebastian Raschka, PhD's avatar" height="36" src="https://substackcdn.com/image/fetch/$s_!CfW_!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F61f4c017-506f-4e9b-a24f-76340dad0309_800x800.jpeg" width="36"&gt;&lt;/img&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;                    &lt;div&gt;                      &lt;div&gt;                        &lt;a href="https://substack.com/@rasbt"&gt;Sebastian Raschka, PhD&lt;/a&gt;&lt;/div&gt;                      &lt;div&gt;                        &lt;div&gt;Sep 29, 2026&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;                &lt;div&gt;                  &lt;div&gt;                    &lt;div&gt;                      &lt;div&gt;            &lt;br /&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;            &lt;div&gt;&lt;/div&gt;            &lt;div&gt;              &lt;div&gt;                &lt;div&gt;                  &lt;div&gt;                    &lt;p&gt;The recently released                      &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;Jev&lt;/a&gt;AI model has been quite a cultural phenomenon in technical communities in the past 2 weeks.&lt;/p&gt;                    &lt;p&gt;While Jev aims to classify things, it’s easy to dismiss Jev as “just a classifier,” and my own view of Jev has evolved quite a bit over the past few days. In particular, my thoughts went from “classifiers used to be my bread &amp;amp; butter; I can easily build this myself” (more on this later) to “wow, this actually works better than I thought.”&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!0Q-x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5481509-dd15-42bf-9ce0-29a292276016_6461x3353.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="jev-intro" height="756" src="https://substackcdn.com/image/fetch/$s_!0Q-x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5481509-dd15-42bf-9ce0-29a292276016_6461x3353.png" title="jev-intro" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 1: Quick overview of the Jev API; more details on that later.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;Sure, the latest state-of-the-art GPT and open-weight LLMs can do the same kinds of classification tasks as Jev, while also being capable of much more general decision-making. But Jev’s advantage is that it can handle those classification tasks much faster and more cheaply.&lt;/p&gt;                    &lt;p&gt;At the other end of the spectrum, for a narrow, well-defined problem, Jev probably won’t classify anything better, faster, or cheaper than a special-purpose classifier. But its selling point is that it is far more general than those task-specific models.&lt;/p&gt;                    &lt;p&gt;So, what is the methodology behind Jev (based on an educated guess), what can it do, and why is it so popular? I aim to answer all of these later in this article. However, I thought starting with a brief history of language models for decision-making would be a great way to begin. And it hopefully helps demystify some of the hype and show what Jev does very well (”Jev is essentially a text classifier,” but “Jev is also not ‘just’ a text classifier.”)&lt;/p&gt;                    &lt;p&gt;                      &lt;em&gt;PS: I am not affiliated with Jev in any way. Also, I am not offered free access to Jev, and this is also not a product endorsement, just a technical article to offer some insights into the history of text classification to help you make sense of the recent hype.&lt;/em&gt;&lt;/p&gt;                    &lt;p&gt;Since this is a long article,                      &lt;strong&gt;I recommend                        &lt;a href="https://magazine.sebastianraschka.com/p/classifier-history-and-jev"&gt;reading it in your browser&lt;/a&gt;,&lt;/strong&gt;where you can access the table of contents menu on the left side.&lt;/p&gt;                    &lt;h2&gt;1. Language modeling and classification in the pre-transformer era                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;                    &lt;p&gt;For completeness, before we put Jev in context (no pun intended), I thought it made the most sense to start chronologically. In this section, I want to take a brief tour of applied text classification via naive Bayes, logistic regression, and the more classic (deep) neural networks before transformer-based models came along.&lt;/p&gt;                    &lt;h3&gt;1.1 Bag-of-words: naive Bayes, logistic regression, and XGBoost                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;                    &lt;p&gt;Back in the day, when I was a grad student 15 years ago, even though recurrent neural networks already existed (more on that later), text classification was usually done with a bag-of-words representation because it was straightforward and could get good results on moderately sized datasets.&lt;/p&gt;                    &lt;p&gt;In short, we can think of the bag-of-words representation as a method that makes free-form text input of different lengths compatible with classic classifiers (naive Bayes, Logistic Regression, SVMs, Random Forest, XGBoost, to name a few), which expect a fixed-size input vector.&lt;/p&gt;                    &lt;p&gt;Popular real-world applications include anything from news article classification to email spam filtering. And yes, allegedly even Gmail’s original spam filter used a Naive Bayes model with a bag-of-words representation.&lt;/p&gt;                    &lt;p&gt;As a side note,                      &lt;a href="https://arxiv.org/abs/1410.5329"&gt;I wrote about this&lt;/a&gt;approach exactly 12 years ago. It was one of the first things I shared on arXiv.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!UgdR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7900489a-1e4c-4997-974b-3b5f89d62465_1050x1318.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="" height="509.62666666666667" src="https://substackcdn.com/image/fetch/$s_!UgdR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7900489a-1e4c-4997-974b-3b5f89d62465_1050x1318.png" width="406"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 2: An                      &lt;a href="https://arxiv.org/abs/1410.5329"&gt;old tutorial&lt;/a&gt;of mine from 2014 that explains naive Bayes classifiers using a bag-of-words model.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;So, what exactly is this bag-of-words representation? It’s a way to convert free-form texts with different lengths, e.g.,&lt;/p&gt;                    &lt;ul&gt;                      &lt;li&gt;                        &lt;p&gt;Training example 1:                          &lt;em&gt;“Zentropa is the most original movie I’ve seen in years. If you like unique thrillers that are influenced by film noir, then this is just the right cure for all of those Hollywood summer blockbusters clogging the theaters these days. Von Trier’s follow-ups like Breaking the Waves have gotten more acclaim, but this is really his best work.”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;                      &lt;li&gt;                        &lt;p&gt;Training example 2:                          &lt;em&gt;“This film is just plain horrible. John Ritter doing pratt falls, 75% of the actors delivering their lines as if they were reading them from cue cards, poor editing, horrible sound mixing”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;                      &lt;li&gt;                        &lt;p&gt;Training example 3:                          &lt;em&gt;“Zentropa has much in common with The Third Man, another noir-like film set among the rubble of postwar Europe.”&lt;/em&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;                    &lt;p&gt;into a fixed-size representation for the aforementioned “classic” classifiers. (The example above is an excerpt from the popular                      &lt;a href="https://ai.stanford.edu/~amaas/data/sentiment/"&gt;IMDb movie review classification dataset&lt;/a&gt;.)&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!WLtd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f041db0-58bb-43db-a07c-ac672b613470_7769x3701.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="bow" height="694" src="https://substackcdn.com/image/fetch/$s_!WLtd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f041db0-58bb-43db-a07c-ac672b613470_7769x3701.png" title="bow" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 3: An illustration of a bag-of-words representation.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;A bag-of-words model starts by building the vocabulary, which consists of all unique words in the training set (optionally, one can get rid of so-called stopwords like “a” and “the”, which are words that carry little to no semantic meaning in most contexts).&lt;/p&gt;                    &lt;p&gt;A bag-of-words representation results in these fixed-size inputs by assigning each word in a vocabulary its own position in a vector. We then count how often each word occurs in a document. For example, if we have a vocabulary of 50,000 unique words, it produces a fixed-size vector with 50,000 entries, regardless of whether the input consists of only ten words or 300k words. Note that most entries are zero because each document contains only a small subset of the vocabulary. (Instead of representing the raw counts, there are also normalization schemes like TF-IDF.)&lt;/p&gt;                    &lt;p&gt;Then, once we have these word frequency vectors, we can train a classifier on a labeled training set, such as emails labeled as spam or non-spam. For example, a logistic regression model would then learn feature weights that correlate certain words (and word counts) with particular labels. For instance, certain words might increase the predicted spam probability, and others may decrease it.&lt;/p&gt;                    &lt;p&gt;This approach is computationally cheap and can work well when particular words provide strong clues about the label. In a simple classification task such as spam classification, this is often enough to get quick, reasonably accurate results.&lt;/p&gt;                    &lt;p&gt;But one of the biggest downsides of this approach is that, because of the nature of the bag-of-words representation, it loses word order. So, for example, “the dog bites the man” and “the man bites the dog” produce identical vectors despite describing different events.&lt;/p&gt;                    &lt;p&gt;(There are some workarounds to preserve some local order by adding word pairs or longer sequences, called n-grams, as features, although this increases the vocabulary size.)&lt;/p&gt;                    &lt;p&gt;Despite the shortcomings, I still think that a bag-of-words has its place in certain low-stakes applications because it’s so cheap, and a bag-of-words representation + logistic regression remains my go-to baseline for every text classification problem, since it’s so easy to implement.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!qNQv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbb03a2-bb84-407b-8e08-f298e72a85f2_3907x2469.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="logreg" height="920" src="https://substackcdn.com/image/fetch/$s_!qNQv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1fbb03a2-bb84-407b-8e08-f298e72a85f2_3907x2469.png" title="logreg" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 4: For those interested in building a simple logistic regression classifier, I have a tutorial up                      &lt;a href="https://github.com/rasbt/machine-learning-book/blob/main/ch08/ch08.ipynb"&gt;here&lt;/a&gt;. This model achieves 89.9% accuracy (on a balanced dataset).&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;h3&gt;1.2 Deep neural networks for text classification                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;                    &lt;p&gt;The aforementioned bag-of-words model would also work with (simple) deep neural networks, like multilayer perceptrons. But the downside still is that we would lose the sentence structure and word order.&lt;/p&gt;                    &lt;p&gt;However, more sophisticated neural network architectures avoid the bag-of-words workaround: convolutional neural networks (CNNs) and recurrent neural networks (RNNs), which can take word embeddings as input.&lt;/p&gt;                    &lt;h4&gt;1.2.1 Word embeddings                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h4&gt;                    &lt;p&gt;First, before feeding the input texts into a model, we have to convert them into a suitable representation. One such representation is bag-of-words. Another is word embedding vectors. The difference is that a bag-of-words vector represents the entire text by counting how often each vocabulary word occurs, while a word embedding represents an individual word as a dense vector of learned numbers.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!2shX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7384f46-af12-44c5-a1b4-5ec0c023094c_7579x3181.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="word-embeddings" height="611" src="https://substackcdn.com/image/fetch/$s_!2shX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7384f46-af12-44c5-a1b4-5ec0c023094c_7579x3181.png" title="word-embeddings" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 5: Illustration of creating word embeddings.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;ul&gt;                      &lt;li&gt;                        &lt;p&gt;Word embeddings work similarly to embedding layers in LLMs, i.e., they convert input tokens into dense vectors. Embeddings can happen outside the model (e.g., two classic, popular methods for learning them are                          &lt;a href="https://arxiv.org/abs/1301.3781"&gt;Word2Vec&lt;/a&gt;and                          &lt;a href="https://aclanthology.org/D14-1162/"&gt;GloVe&lt;/a&gt;), or the embedding layer can be part of the neural network architecture itself and be learned and tuned during model training.&lt;/p&gt;                        &lt;p&gt;These classic embeddings are context-independent at lookup time. The word “bank”, for example, gets the same vector in “river bank” and “bank account” (as you may know, context can be handled via concepts like attention).&lt;/p&gt;                        &lt;p&gt;For additional resources on embeddings, you may find the following ones helpful:&lt;/p&gt;                        &lt;ul&gt;                          &lt;li&gt;                            &lt;p&gt;                              &lt;a href="https://github.com/rasbt/LLMs-from-scratch/blob/main/ch02/01_main-chapter-code/ch02.ipynb"&gt;Chapter 2: Working with Text Data&lt;/a&gt;(this is an LLM chapter but should give you the gist of embedding words or tokens; in LLM tokenizers, we split words into subword tokens; in Word2Vec, 1 word is usually 1 token.)&lt;/p&gt;&lt;/li&gt;                          &lt;li&gt;                            &lt;p&gt;                              &lt;a href="https://github.com/rasbt/LLMs-from-scratch/blob/main/ch02/03_bonus_embedding-vs-matmul/embeddings-and-linear-layers.ipynb"&gt;Understanding the Difference Between Embedding Layers and Linear Layers&lt;/a&gt;(this is an illustration that shows when embedding vectors are mathematically equivalent to Linear layers and matrix multiplications.)&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;&lt;/ul&gt;                    &lt;h4&gt;1.2.2 Recurrent neural networks (RNNs)                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h4&gt;                    &lt;p&gt;Since many of you are probably familiar with recurrent neural networks (RNNs), I will keep this section short. RNNs are a classic go-to neural network architecture for natural language processing, and popular  variants go back to the 1980s and early 1990s. Transformers, which were introduced in 2017 (and using the attention mechanisms that were first introduced in RNNs; see my                      &lt;a href="https://magazine.sebastianraschka.com/p/understanding-large-language-models"&gt;Understanding Large Language Models&lt;/a&gt;for a brief timeline), then gradually replaced them in many NLP applications.&lt;/p&gt;                    &lt;p&gt;RNNs read a sequence (like text) one word at a time. At each step, they combine the current word embedding (discussed in the previous section) with a hidden state from the previous step. We can think of the hidden state as a fixed-size vector that summarizes the text processed so far, so this makes word order matter, because rearranging the words changes the sequence of state updates.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!k4d9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c0da28-827d-4a59-9fed-cb283e57ad1b_7776x4745.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="rnn-classify" height="888" src="https://substackcdn.com/image/fetch/$s_!k4d9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c0da28-827d-4a59-9fed-cb283e57ad1b_7776x4745.png" title="rnn-classify" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 6: Illustration of an RNN classifier.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;Note that the figure above shows the RNN in the unrolled representation. I.e., the RNN reuses the same layer stack for each input, hence the term “recurrent”. And since it’s “recurrent”, the input text can have an arbitrary length. The figure below illustrates the “recurrence” with the unrolled representation side by side. Note that both show the identical architecture, it’s just a different visualization.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!JOy3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc3e8a8d-8e3b-4989-8b42-564cf84c80b6_7675x3257.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="rolled-rnn" height="618" src="https://substackcdn.com/image/fetch/$s_!JOy3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc3e8a8d-8e3b-4989-8b42-564cf84c80b6_7675x3257.png" title="rolled-rnn" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 7: Rolled and unrolled illustration of the same RNN.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;RNNs were notoriously hard to train, and there are important improvements to RNNs, like                      &lt;a href="https://www.bioinf.jku.at/publications/older/2604.pdf"&gt;Long short-term memory&lt;/a&gt;(LSTM) networks, introduced in 1997, and                      &lt;a href="https://aclanthology.org/D14-1179/"&gt;gated recurrent units&lt;/a&gt;(GRUs), introduced in 2014, which use learned gates to control how information is retained and updated. (There is also the more recent                      &lt;a href="https://proceedings.neurips.cc/paper_files/paper/2024/hash/c2ce2f2701c10a2b2f2ea0bfa43cfaa3-Abstract-Conference.html"&gt;xLSTM: Extended long short-term memory&lt;/a&gt;, introduced in 2024).&lt;/p&gt;                    &lt;p&gt;Also, state-space models are inspired by this idea of a fixed-size hidden state updated sequentially, which is cheaper than transformer attention. However, the bottleneck is still how much information the hidden state can retain, and it still has to be processed sequentially. (Fun fact: attention was first developed for RNNs before the transformer architecture came along, but it’s a story for another time; I’ve written about it in my                      &lt;a href="https://magazine.sebastianraschka.com/p/understanding-large-language-models"&gt;Understanding Large Language Models&lt;/a&gt;article.)&lt;/p&gt;                    &lt;p&gt;The bottom line is that RNNs can be used to train text classifiers. Coming back to the IMDb movie review dataset, the bag-of-words classifier with logistic regression achieved about 89.9% accuracy (on a balanced dataset), while an LSTM RNN achieved only 85.66% accuracy. Yes, RNNs can be harder to train (stay tuned for the ULMFiT method below, which trains an RNN with much higher accuracy).&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!PvQO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbb8bd45-04a6-47c5-bbdb-b87fe171b32b_3866x2127.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="rnn-imdb" height="322.3804945054945" src="https://substackcdn.com/image/fetch/$s_!PvQO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbb8bd45-04a6-47c5-bbdb-b87fe171b32b_3866x2127.png" title="rnn-imdb" width="586"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 8: A simple RNN with LSTM                      &lt;a href="https://github.com/rasbt/machine-learning-book/blob/main/ch15/ch15_part2.ipynb"&gt;tutorial&lt;/a&gt;. This model achieves 85.66% accuracy (on a balanced dataset); the higher training accuracy indicates substantial overfitting.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;Note that this RNN was trained from scratch. A better approach is to pre-train the model on a larger dataset first, then fine-tune it on this target dataset (classically, we call this approach “transfer learning”).&lt;/p&gt;                    &lt;p&gt;In the natural language processing domain, one of the most influential papers proposing this approach is                      &lt;a href="https://arxiv.org/abs/1801.06146"&gt;ULMFiT&lt;/a&gt;(2018), which achieved an impressive 95.4% test accuracy on IMDb.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!HZOG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e8aeef2-95e4-415a-8f5f-93895d1f2eb2_4470x2477.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="ulmfit" height="807" src="https://substackcdn.com/image/fetch/$s_!HZOG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e8aeef2-95e4-415a-8f5f-93895d1f2eb2_4470x2477.png" title="ulmfit" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 9: Annotated figure from the                      &lt;a href="https://arxiv.org/abs/1801.06146"&gt;ULMFiT&lt;/a&gt;paper.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;h4&gt;1.2.3 Convolutional neural networks (CNNs)                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h4&gt;                    &lt;p&gt;You probably know convolutional neural networks (CNNs) from their use in computer vision. However, it is also possible, although historically less common, to use them for text.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!ZK8R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbacd70f8-4159-41b2-a812-9ab65de29f2b_7288x2635.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="cnn-vision" height="526" src="https://substackcdn.com/image/fetch/$s_!ZK8R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbacd70f8-4159-41b2-a812-9ab65de29f2b_7288x2635.png" title="cnn-vision" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 10: Illustration of a convolutional neural network (CNN) to classify images.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;As shown in the image classification example in the figure above, CNNs apply learned filters to image patches (windows). Similarly, in the natural language domain, we can apply learned filters to windows of adjacent word embeddings.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!zRgr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b710ad1-6fad-47a7-9b79-192b61097194_7935x7279.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="cnn-all" height="1336" src="https://substackcdn.com/image/fetch/$s_!zRgr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b710ad1-6fad-47a7-9b79-192b61097194_7935x7279.png" title="cnn-all" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 11: A CNN for text classification, step by step. Only 1 filter (channel) for simplicity.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;As illustrated in the figure above, a convolutional filter with a window size of three uses the same weights for each adjacent group of three words. Then, as it slides over the inputs, it moves the filter one word at a time across “the movie had surprisingly good acting”. So, this gives four windows (ignoring padding for simplicity):&lt;/p&gt;                    &lt;ul&gt;                      &lt;li&gt;                        &lt;p&gt;1: “the movie had”&lt;/p&gt;&lt;/li&gt;                      &lt;li&gt;                        &lt;p&gt;2: “movie had surprisingly”&lt;/p&gt;&lt;/li&gt;                      &lt;li&gt;                        &lt;p&gt;3: “had surprisingly good”&lt;/p&gt;&lt;/li&gt;                      &lt;li&gt;                        &lt;p&gt;4: “surprisingly good acting”&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;                    &lt;p&gt;So, in the last layer, before the classification head, we can either flatten or global max pool the results before connecting it to the classification head. While flattening preserves all information, it would produce differently sized vectors depending on the text input length. (E.g., if “the movie had surprisingly good acting” were longer, we would have longer feature maps.) So, to make it input-length agnostic, global max-pooling would be a better option here.&lt;/p&gt;                    &lt;p&gt;In short, we can visualize the text CNN as shown below.&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!7qXx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88ad7444-1759-4966-a53e-4744f13ff889_7748x2519.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="cnn-summary" height="473" src="https://substackcdn.com/image/fetch/$s_!7qXx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88ad7444-1759-4966-a53e-4744f13ff889_7748x2519.png" title="cnn-summary" width="1456"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 12: CNN for text classification with multiple filters (channels).&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;p&gt;Also, what’s nice is that the filter computations across positions can run in parallel, which avoids the step-by-step dependency of an RNN.&lt;/p&gt;                    &lt;p&gt;For a quick comparison, on the aforementioned IMDb dataset, my experiments show that such a CNN gets about 90.07% accuracy (but note that this is highly architecture-dependent; for example, you may know from computer vision contexts that accuracies can vary widely). (E.g., the good old AlexNet had a ~62.5% top-1 accuracy on ImageNet, and a ConvNeXt V2-H gets 88.9%.)&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://substackcdn.com/image/fetch/$s_!CSes!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91765c78-a4d8-48bf-ade0-10b515441b8a_4442x2189.png" target="_blank"&gt;                        &lt;div&gt;                          &lt;img alt="cnn-imdb" height="299.8241758241758" src="https://substackcdn.com/image/fetch/$s_!CSes!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91765c78-a4d8-48bf-ade0-10b515441b8a_4442x2189.png" title="cnn-imdb" width="608"&gt;&lt;/img&gt;                          &lt;div&gt;                            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 13: A CNN trained to classify IMDb movie reviews; it gets 90.07% accuracy on this balanced dataset. You can find the source code                      &lt;a href="https://github.com/rasbt/deeplearning-models/blob/master/pytorch_ipynb/cnn-nlp/cnn_imdb.ipynb"&gt;here&lt;/a&gt;.&lt;/div&gt;                    &lt;p&gt;&lt;/p&gt;                    &lt;h2&gt;2. Transformers                      &lt;div&gt;                        &lt;div&gt;                          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;                    &lt;p&gt;In 2017, the original transformer architecture was introduced in the                      &lt;a href="https://arxiv.org/abs/1706.03762"&gt;Attention Is All You Need&lt;/a&gt;paper. Since this topic has been covered so extensively (by me and others), I will focus on the classification-relevant aspects. But for those interested in the attention mechanism and other architecture details, please see my related articles:&lt;/p&gt;                    &lt;div&gt;                      &lt;a href="https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison" rel="noopener" target="_blank"&gt;                        &lt;div&gt;                          &lt;div&gt;                            &lt;img alt="The Big LLM Architecture Comparison" height="140" src="https://substackcdn.com/image/fetch/$s_!LmVE!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45c50202-0e8b-4e64-8296-4e2ccf4cb287_1756x1227.png" width="140"&gt;&lt;/img&gt;&lt;/div&gt;                          &lt;div&gt;                            &lt;h4&gt;The Big LLM Architecture Comparison&lt;/h4&gt;                            &lt;div&gt;                              &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;                      &lt;a href="https://substack.com/profile/27393275-sebastian-raschka-phd"&gt;Sebastian Raschka, PhD&lt;/a&gt;&lt;/div&gt;                    &lt;div&gt;·&lt;/div&gt;                    &lt;div&gt;July 19, 2025&lt;/div&gt;&lt;/div&gt;                  &lt;div&gt;                    &lt;a href="https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison"&gt;                      &lt;div&gt;Read full story&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;            &lt;div&gt;              &lt;a href="https://magazine.sebastianraschka.com/p/visual-attention-variants" rel="noopener" target="_blank"&gt;                &lt;div&gt;                  &lt;div&gt;                    &lt;img alt="A Visual Guide to Attention Variants in Modern LLMs" height="140" src="https://substackcdn.com/image/fetch/$s_!8IKa!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51d52b9a-e820-45d6-8135-f94496ec1745_1600x900.png" width="140"&gt;&lt;/img&gt;&lt;/div&gt;                  &lt;div&gt;                    &lt;h4&gt;A Visual Guide to Attention Variants in Modern LLMs&lt;/h4&gt;                    &lt;div&gt;                      &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;              &lt;a href="https://substack.com/profile/27393275-sebastian-raschka-phd"&gt;Sebastian Raschka, PhD&lt;/a&gt;&lt;/div&gt;            &lt;div&gt;·&lt;/div&gt;            &lt;div&gt;Mar 22&lt;/div&gt;&lt;/div&gt;          &lt;div&gt;            &lt;a href="https://magazine.sebastianraschka.com/p/visual-attention-variants"&gt;              &lt;div&gt;Read full story&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;The key point is that the original transformer architecture was an encoder-decoder setup used for language translation, but it can be easily adapted for text classification tasks, as I’ll illustrate in the following sections.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!y9UZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd94764e4-62ed-4c95-b468-80fb6f9a6621_6217x7963.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="attention-is-all-you-need" height="402.2046703296703" src="https://substackcdn.com/image/fetch/$s_!y9UZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd94764e4-62ed-4c95-b468-80fb6f9a6621_6217x7963.png" title="attention-is-all-you-need" width="314"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 14: The original transformer architecture from      &lt;a href="https://arxiv.org/abs/1706.03762"&gt;Attention Is All You Need&lt;/a&gt;.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;h3&gt;2.1 Encoder-style language models      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;I remember all too well the first years after the original transformer architecture release. They were pretty much defined by the rivalry between two different approaches:&lt;/p&gt;    &lt;ol&gt;      &lt;li&gt;        &lt;p&gt;Encoder-style models like BERT, primarily developed by Google;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;Decoder-style models like GPT, primarily developed by OpenAI.&lt;/p&gt;&lt;/li&gt;&lt;/ol&gt;    &lt;p&gt;Encoder-style models were natural text classifiers, whereas GPT models could do zero- and few-shot classification as an emergent property, but their strength was more in generative tasks.&lt;/p&gt;    &lt;p&gt;But let’s start with encoder-style models and how to fine-tune and use them to classify text.&lt;/p&gt;    &lt;p&gt;One universal aspect of using transformers, whether encoder- or decoder-style, is that we work with models pre-trained on large text corpora, and, in addition to using them as zero- or few-shot classifiers, we can fine-tune them on the target dataset (similar to ULMFiT, as mentioned in the RNN section earlier).&lt;/p&gt;    &lt;p&gt;As shown in the figure excerpt from the      &lt;a href="https://arxiv.org/abs/1810.04805"&gt;BERT paper&lt;/a&gt;(2018) below, BERT-/encoder-style models have a classification token at the first position that we can conveniently fine-tune for this.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!zrSS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59456abf-cf68-4d6e-9c92-73ad3e593665_4982x3213.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="bert-annotate" height="405.00824175824175" src="https://substackcdn.com/image/fetch/$s_!zrSS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59456abf-cf68-4d6e-9c92-73ad3e593665_4982x3213.png" title="bert-annotate" width="628"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 15: Annotated figure from the original      &lt;a href="https://arxiv.org/abs/1810.04805"&gt;BERT paper&lt;/a&gt;.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;While BERT models are not nearly as popular as autoregressive GPT-style transformers, thankfully some people still update and modernize them occasionally. One recent example that is often my go-to for classification tasks is the 2024      &lt;a href="https://arxiv.org/abs/2412.13663"&gt;ModernBERT&lt;/a&gt;model.&lt;/p&gt;    &lt;p&gt;As shown below, on the IMDb movie reviews, ModernBERT gets approximately 95% accuracy with very little fine-tuning effort.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!GqmY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4cc75fe-608d-47fc-95aa-14f3f2b58b83_4147x2928.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="bert-imdb" height="321.95604395604397" src="https://substackcdn.com/image/fetch/$s_!GqmY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4cc75fe-608d-47fc-95aa-14f3f2b58b83_4147x2928.png" title="bert-imdb" width="456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 16: Different LLMs fine-tuned to classify IMDb movie reviews; you can find the source code      &lt;a href="https://github.com/rasbt/LLMs-from-scratch/tree/main/ch06/03_bonus_imdb-classification"&gt;here&lt;/a&gt;. (I did very little hyperparameter tuning, and this can potentially be improved by 1-2% further; however, one of the points here is also that one can get really good results with pretty minimal effort.)&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;h3&gt;2.2 Decoder-style LLMs      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;LLMs like GPT are decoder-style, autoregressive transformers that are the center of all attention (no pun intended) for generating text and code.&lt;/p&gt;    &lt;p&gt;However, as I explained in chapter 6 of my      &lt;a href="https://amzn.to/4fqvn0D"&gt;Build A Large Language Model (From Scratch)&lt;/a&gt;book, as a gentle introduction to fine-tuning (before covering instruction fine-tuning), we can also repurpose these for text classification.&lt;/p&gt;    &lt;p&gt;Sure, we can also prompt an LLM directly, as shown in the screenshot below.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!Rh-F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be7e4aa-5436-46cf-8e85-809f776f39e1_4920x3189.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="prompt-gpt" height="382.5274725274725" src="https://substackcdn.com/image/fetch/$s_!Rh-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be7e4aa-5436-46cf-8e85-809f776f39e1_4920x3189.png" title="prompt-gpt" width="590"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 17: Prompting an LLM to classify a movie review.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;However, if we want structured outputs and we have a specific target domain in mind, this is unnecessarily brittle and inefficient.&lt;/p&gt;    &lt;p&gt;Instead, we can replace the output layer with a leaner classification head, as illustrated below:&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!T3RZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9065aa90-305d-4373-bf45-73d82ec86768_1943x1931.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="gpt-head" height="524.7362637362637" src="https://substackcdn.com/image/fetch/$s_!T3RZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9065aa90-305d-4373-bf45-73d82ec86768_1943x1931.png" title="gpt-head" width="528"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 18: Swapping the output layer of a GPT-style model with a leaner classification head.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Now, fine-tuning has a few caveats. For instance, because of the autoregressive attention mask, we have to be careful about how we design fine-tuning so that the classification token has information about all other tokens in the sequence.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!5cCi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bc5509-085e-4b93-b557-566a5a05f0ff_2940x2021.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="attention-masks" height="361.625" src="https://substackcdn.com/image/fetch/$s_!5cCi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bc5509-085e-4b93-b557-566a5a05f0ff_2940x2021.png" title="attention-masks" width="526"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 19: Attention masks in BERT and GPT.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;If you are interested in further technical details, please see my recent end-to-end      &lt;a href="https://magazine.sebastianraschka.com/p/ai-detector-from-scratch"&gt;Building an AI Text Detector From Scratch&lt;/a&gt;article:&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;        &lt;a href="https://magazine.sebastianraschka.com/p/ai-detector-from-scratch" rel="noopener" target="_blank"&gt;          &lt;h2&gt;Building an AI Text Detector From Scratch&lt;/h2&gt;&lt;/a&gt;        &lt;div&gt;          &lt;div&gt;            &lt;a href="https://substack.com/profile/27393275-sebastian-raschka-phd"&gt;Sebastian Raschka, PhD&lt;/a&gt;&lt;/div&gt;          &lt;div&gt;·&lt;/div&gt;          &lt;div&gt;Aug 15&lt;/div&gt;&lt;/div&gt;        &lt;div&gt;          &lt;a href="https://magazine.sebastianraschka.com/p/ai-detector-from-scratch" rel="noopener" target="_blank"&gt;            &lt;img alt="Building an AI Text Detector From Scratch" height="650" src="https://substackcdn.com/image/fetch/$s_!Om23!,w_1300,h_650,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F875b5076-19a9-4ffe-91d4-cafa156f650a_5343x5696.png" width="1300"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/div&gt;        &lt;div&gt;          &lt;p&gt;Substack recently launched its AI detector feature in the UI, which is super interesting.&lt;/p&gt;&lt;/div&gt;        &lt;div&gt;          &lt;a href="https://magazine.sebastianraschka.com/p/ai-detector-from-scratch"&gt;            &lt;div&gt;Read full story&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;Overall, though, the advantage of using GPT-style LLMs over e.g., BERT variants is that there are so many modern open-weight architectures out there to adopt. Anything between the small Qwen 3 0.6B models to the latest Kimi, GLM, or DeepSeek models. (Of course, using such &amp;gt;1B parameter models for classification could be a bit overkill from an efficiency perspective, but hey, it’s possible.)&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!z4j3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41880c4-c2e4-4deb-b97b-13d82b47c8c7_4164x2928.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="gpt-classify" height="372.74725274725273" src="https://substackcdn.com/image/fetch/$s_!z4j3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb41880c4-c2e4-4deb-b97b-13d82b47c8c7_4164x2928.png" title="gpt-classify" width="530"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;    &lt;p&gt;Figure 20: On the aforementioned IMDb movie review dataset, a relatively small GPT-2 124M model gets approximately 92% accuracy; a larger and newer LLM (like a recent Qwen3 variant) would likely perform better.&lt;/p&gt;    &lt;h3&gt;2.3 Encoder-decoder style architectures      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;While the original transformer architecture (an encoder-decoder) was split into two paradigms, encoder-style models like BERT and decoder-style LLMs like GPT, there were also efforts to use encoder-decoder-style variants, with the most recent variant (somewhat unexpectedly) being      &lt;a href="https://sebastianraschka.com/llm-architecture-gallery/#card-deepseek-v4-1-flash"&gt;DeepSeek V4.1 Flash&lt;/a&gt;. (However, in this case, it’s a causal encoder not a bidirectional one as in T5.)&lt;/p&gt;    &lt;p&gt;Keeping the focus on encoder-decoder architectures for classification, probably the most prominent candidate is Google’s 2019      &lt;a href="https://arxiv.org/abs/1910.10683"&gt;T5 (Text-to-Text Transfer Transformer)&lt;/a&gt;.&lt;/p&gt;    &lt;p&gt;Architecturally, the differences between T5 and the original transformer include updates to the architecture (as summarized below), as well as changes to training. The original transformer was trained to do language translation (in a supervised fashion); T5 uses pre-training on unlabeled text with span corruption, where the encoder receives text with missing spans and the decoder generates those missing spans.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!EeEQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb15fdb01-219a-4373-8b17-a15a9956436a_6256x4375.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="t5" height="1018" src="https://substackcdn.com/image/fetch/$s_!EeEQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb15fdb01-219a-4373-8b17-a15a9956436a_6256x4375.png" title="t5" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 21: T5 architecture changes to the original transformer architecture. (The architecture drawing is taken from      &lt;a href="https://arxiv.org/abs/1706.03762"&gt;Attention Is All You Need&lt;/a&gt;.)&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Now, we can use T5 similar to the other RNN, CNN, and transformer approaches above, by adding a classification head.&lt;/p&gt;    &lt;p&gt;Additionally, we can also (train to) have the decoder output the class label prediction (like “positive” or “negative” in the IMDb movie review dataset case), similar to regular LLMs. Let’s call this approach text-to-text classification.&lt;/p&gt;    &lt;p&gt;While GPT-style LLMs are trained on massive amounts of text, they usually perform text-to-text classification well out of the box, as shown earlier (the figure below is inserted here again for convenience).&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!Rh-F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be7e4aa-5436-46cf-8e85-809f776f39e1_4920x3189.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="prompt-gpt" height="414.94505494505495" src="https://substackcdn.com/image/fetch/$s_!Rh-F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be7e4aa-5436-46cf-8e85-809f776f39e1_4920x3189.png" title="prompt-gpt" width="640"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 22: Text-to-text classification with a GPT model.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;However, in the case of T5, it is common to further fine-tune the decoder to do well on these types of tasks in a given target domain.&lt;/p&gt;    &lt;p&gt;For T5, both approaches work. Here’s a summary:&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!bg1N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2a84f80-0975-4ddc-aee6-d3579f50edc5_5046x1890.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="Classification head vs text-to-text approaches" height="545" src="https://substackcdn.com/image/fetch/$s_!bg1N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2a84f80-0975-4ddc-aee6-d3579f50edc5_5046x1890.png" title="Classification head vs text-to-text approaches" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 23: Classification head vs text-to-text approaches.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;h2&gt;3. Jev overview      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;    &lt;p&gt;So far, we have seen that there are plenty of approaches to text classification, from traditional methods like logistic regression and naive Bayes with bag-of-words representations to prompting the latest frontier LLMs like GPT-6 or fine-tuning any open-weight LLM with a classification head (or text-to-text classification).&lt;/p&gt;    &lt;p&gt;At first, I (almost) dismissed it as “just a classifier,” something that I build routinely for classification tasks with natural language inputs (e.g., see my      &lt;a href="https://magazine.sebastianraschka.com/p/ai-detector-from-scratch"&gt;AI Detector article&lt;/a&gt;as a recent public example).&lt;/p&gt;    &lt;h3&gt;3.1 Jev vs existing text-to-text classification      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Before continuing this discussion, though, let’s start with a quick Jev overview.      &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;Jev is a new model released by TypeSafe AI&lt;/a&gt;, which just came out of stealth a few weeks ago and got a lot of attention (at first, it seemed a bit bizarre because it looks just like a classifier).&lt;/p&gt;    &lt;p&gt;Jev is a proprietary model (although the release was followed by a huge number of quick open-source clones, but we will get to this later) that is relatively cheap to use and claims to be on par with GPT-5.6 Luna (for decision-making) while being orders of magnitude faster and cheaper:&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!2oFN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d2d71d-7704-433d-9d4d-ac8a87ef3b8a_5584x3084.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-bench" height="804" src="https://substackcdn.com/image/fetch/$s_!2oFN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13d2d71d-7704-433d-9d4d-ac8a87ef3b8a_5584x3084.png" title="jev-bench" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 24: Benchmark from      &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;TypeSafe AI blog&lt;/a&gt;.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Here, regarding GPT-Luna, think of this as the text-to-text classification approach discussed earlier.&lt;/p&gt;    &lt;p&gt;So why all this hype? I think it’s partly because of the nice API and that it performs so well on all kinds of tasks, so it doesn’t require custom fine-tuning.&lt;/p&gt;    &lt;p&gt;For example, I can use it to categorize support tickets/emails (I made a video version to show the response time):&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Or, I can have the same model play Tetris in real-time (here, I am using the Choice API):&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Maybe the best way to succinctly explain it (before getting into technical details) is as follows: Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.&lt;/p&gt;    &lt;p&gt;(As of this writing, there’s not much known about the architecture and exact training algorithm, except that one of the founders said it was trained via “Reinforcement Learning for Calibrated Decisions”, but more on that later.)&lt;/p&gt;    &lt;h3&gt;3.2 The Jev API      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Over the past few years, we’ve gotten used to throwing bigger, better, and more expensive GPT-style LLMs at all kinds of problems, and for targeted decision-making or classification tasks, a cheap &amp;amp; fast approach like Jev may feel refreshing to most. Especially for one-off tasks where collecting training data and fine-tuning a custom ModernBERT sounds too tedious, we might just throw a Luna-like model at it.&lt;/p&gt;    &lt;p&gt;But on top of being popular for its versatility (that is, performing well on different target domains out of the box), Jev also has a relatively nice API which we can use via curl in the terminal or via its Python API.&lt;/p&gt;    &lt;p&gt;In short, there are three main API types illustrated below. Let’s start with the Choice API, which is convenient for multi-class classification.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!SzWL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac35084-507b-4f08-856c-dd449f43ebf2_6461x3353.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-choice" height="756" src="https://substackcdn.com/image/fetch/$s_!SzWL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ac35084-507b-4f08-856c-dd449f43ebf2_6461x3353.png" title="jev-choice" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 25: Jev’s Choice API. Use this for multi-class classification.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Next is the Noul API, which is simpler and assigns a “yes” probability to a question.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!8qtX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb63ec307-8be5-48c5-a0f3-21dbe461c78d_6461x3353.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-noul" height="756" src="https://substackcdn.com/image/fetch/$s_!8qtX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb63ec307-8be5-48c5-a0f3-21dbe461c78d_6461x3353.png" title="jev-noul" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 26: Jev’s Noul API. Use this for binary classification or multi-label classification (with multiple Noul questions).&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Lastly, the Score API assigns a score based on a rubric level (in the example below, 0, 1, 2).&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!4RFc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3eb167e-451f-4bb8-8c8d-4da0a781ab5d_6461x3354.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-score" height="756" src="https://substackcdn.com/image/fetch/$s_!4RFc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3eb167e-451f-4bb8-8c8d-4da0a781ab5d_6461x3354.png" title="jev-score" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 27: Jev’s Score API. Use this for ordinal classification.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;To sum it up, the different APIs and use cases are as follows.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!QQ6G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe962a588-c3d3-453a-9e1a-abfccfd4e7b7_4455x1747.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-api-table" height="571" src="https://substackcdn.com/image/fetch/$s_!QQ6G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe962a588-c3d3-453a-9e1a-abfccfd4e7b7_4455x1747.png" title="jev-api-table" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 28: Jev API cheat sheet.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;h3&gt;3.3 Jev classifying IMDb      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;To make these APIs more concrete, and to answer the question of how well Jev might do on the aforementioned IMDb dataset, we can run it with either the Choice or Noul API.&lt;/p&gt;    &lt;p&gt;Let’s start with the Choice API first. Each review would be formatted as follows:&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;        &lt;pre&gt;          &lt;code&gt;export TYPESAFE_API_KEY=&amp;quot;YOUR_API_KEY&amp;quot;

curl -sS https://api.typesafe.ai/v1/systemone \
  -H &amp;quot;Authorization: Bearer $TYPESAFE_API_KEY&amp;quot; \
  -H &amp;quot;Content-Type: application/json&amp;quot; \
  -d &amp;apos;{
    &amp;quot;model&amp;quot;: &amp;quot;jev-1.13.0&amp;quot;,
    &amp;quot;state&amp;quot;: &amp;quot;The acting was excellent and the story kept me engaged throughout. I would happily watch this movie again.&amp;quot;,
    &amp;quot;questions&amp;quot;: {
      &amp;quot;sentiment&amp;quot;: {
        &amp;quot;type&amp;quot;: &amp;quot;choice&amp;quot;,
        &amp;quot;instructions&amp;quot;: &amp;quot;What is the overall sentiment of this movie review?&amp;quot;,
        &amp;quot;criteria&amp;quot;: {
          &amp;quot;negative&amp;quot;: &amp;quot;An overall unfavorable opinion of the movie&amp;quot;,
          &amp;quot;positive&amp;quot;: &amp;quot;An overall favorable opinion of the movie&amp;quot;
        }
      }
    }
  }&amp;apos;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;This Choice API is convenient if we want to use explicit labels for binary (here) or multi-class problems.&lt;/p&gt;    &lt;p&gt;An actual response might look as follows:&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;        &lt;pre&gt;          &lt;code&gt;{
  &amp;quot;model&amp;quot;: &amp;quot;jev-1.13.0&amp;quot;,
  &amp;quot;answers&amp;quot;: {
    &amp;quot;sentiment&amp;quot;: {
      &amp;quot;type&amp;quot;: &amp;quot;choice&amp;quot;,
      &amp;quot;choice&amp;quot;: &amp;quot;positive&amp;quot;,
      &amp;quot;confidence&amp;quot;: 1.0,
      &amp;quot;probabilities&amp;quot;: {
        &amp;quot;negative&amp;quot;: 0.0,
        &amp;quot;positive&amp;quot;: 1.0
      }
    }
  },
  &amp;quot;usage&amp;quot;: {
    &amp;quot;input_tokens&amp;quot;: 342,
    &amp;quot;output_tokens&amp;quot;: 32
  }
}&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;Along with the      &lt;code&gt;&amp;quot;choice&amp;quot;: &amp;quot;positive&amp;quot;&lt;/code&gt;label, we also get the confidence for this prediction, which can be really useful in real-world applications, as well as the probabilities for each class. (The confidence field summarizes how concentrated the probability distribution is. It is different from the probability assigned to the winning class.) One of the selling points of Jev is that, according to the documentation, these are well-calibrated (but more on calibration later).&lt;/p&gt;    &lt;p&gt;Alternatively, we can also use the Noul API here to classify the reviews. Using the same movie review classification context, the format would be as follows:&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;        &lt;pre&gt;          &lt;code&gt;curl -sS https://api.typesafe.ai/v1/systemone \
  -H &amp;quot;Authorization: Bearer $TYPESAFE_API_KEY&amp;quot; \
  -H &amp;quot;Content-Type: application/json&amp;quot; \
  -d &amp;apos;{
    &amp;quot;model&amp;quot;: &amp;quot;jev-1.13.0&amp;quot;,
    &amp;quot;state&amp;quot;: &amp;quot;The acting was excellent and the story kept me engaged throughout. I would happily watch this movie again.&amp;quot;,
    &amp;quot;questions&amp;quot;: {
      &amp;quot;is_positive&amp;quot;: {
        &amp;quot;type&amp;quot;: &amp;quot;noul&amp;quot;,
        &amp;quot;instructions&amp;quot;: &amp;quot;Does this review express an overall positive opinion of the movie?&amp;quot;,
        &amp;quot;criteria&amp;quot;: {
          &amp;quot;true&amp;quot;: &amp;quot;An overall favorable opinion of the movie&amp;quot;,
          &amp;quot;false&amp;quot;: &amp;quot;An overall unfavorable opinion of the movie&amp;quot;
        }
      }
    }
  }&amp;apos;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;And the answer via the Noul API is:&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;        &lt;pre&gt;          &lt;code&gt;{
  &amp;quot;model&amp;quot;: &amp;quot;jev-1.13.0&amp;quot;,
  &amp;quot;answers&amp;quot;: {
    &amp;quot;is_positive&amp;quot;: {
      &amp;quot;type&amp;quot;: &amp;quot;noul&amp;quot;,
      &amp;quot;noul&amp;quot;: 0.98
    }
  },
  &amp;quot;usage&amp;quot;: {
    &amp;quot;input_tokens&amp;quot;: 328,
    &amp;quot;output_tokens&amp;quot;: 21
  }
}&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;(Interestingly, it gives 0.98 instead of 1.0 for the positive class, even though it’s the same text.)&lt;/p&gt;    &lt;p&gt;We can also apply the Noul API to multiple classes by asking one question per class, for example:&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;“Is this article about finance?”&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;“Is this article about politics?”&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;“Is this article about technology?”&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;Here, each question (independently from each other) returns a probability, and these probabilities don’t have to sum to 1. So, for a multi-class classification where each text can have multiple labels, we may prefer Noul, and for classification problems with only one final answer, Choice adds more convenience.&lt;/p&gt;    &lt;p&gt;Now, running Jev with either Choice or Noul on the 25,000 movie reviews on the test set gave the following results:&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;          &lt;strong&gt;Choice:&lt;/strong&gt;&lt;/p&gt;        &lt;ul&gt;          &lt;li&gt;            &lt;p&gt;Accuracy 96.47% (24,117 correct)&lt;/p&gt;&lt;/li&gt;          &lt;li&gt;            &lt;p&gt;Total runtime: 22 minutes 24 seconds&lt;/p&gt;&lt;/li&gt;          &lt;li&gt;            &lt;p&gt;15,456,663 input tokens, $0.6492 total cost&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;          &lt;strong&gt;Noul:&lt;/strong&gt;&lt;/p&gt;        &lt;ul&gt;          &lt;li&gt;            &lt;p&gt;Accuracy 96.20% (24,050 correct)&lt;/p&gt;&lt;/li&gt;          &lt;li&gt;            &lt;p&gt;Total runtime: 23 minutes 3 seconds&lt;/p&gt;&lt;/li&gt;          &lt;li&gt;            &lt;p&gt;15,106,663 input tokens, $0.6345 total cost&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;Note that the small difference in performance and runtime between Choice and Noul could be due to random fluctuation since, similar to other LLMs, the runs are not fully deterministic.&lt;/p&gt;    &lt;p&gt;For instance, I ran the Choice API a second time over the same test set and got slightly different results, as shown below.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!ZBx8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F477155d8-c954-471a-aa18-f456ad16b221_4445x682.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-repeated" height="223" src="https://substackcdn.com/image/fetch/$s_!ZBx8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F477155d8-c954-471a-aa18-f456ad16b221_4445x682.png" title="jev-repeated" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 29: Repeated Jev runs.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;(The non-determinism when serving LLMs and other models at scale is likely due to batch-dependent GPU kernel execution changing the order of floating-point operations, as laid out in a nice      &lt;a href="https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/"&gt;blog post&lt;/a&gt;by Horace He last year.)&lt;/p&gt;    &lt;p&gt;Anyways, the results look quite good overall (caveat: we don’t know if the IMDb test set was part of the training set).&lt;/p&gt;    &lt;p&gt;For reference, a ModernBERT model took&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;23 min to fine-tune on the training set;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;7 min to evaluate on the test set.&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;For comparison, the best ModernBERT model has similar accuracy, as shown below (I expect you can still get 1-2% higher accuracy with additional hyperparameter tuning). Note that this was run on a DGX Spark, and it’s possible to get faster inference performance with quantization and faster hardware.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!mnNk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9933df13-7e9b-4144-86f3-b53781791e85_4776x2169.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="bert-vs-jev" height="661" src="https://substackcdn.com/image/fetch/$s_!mnNk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9933df13-7e9b-4144-86f3-b53781791e85_4776x2169.png" title="bert-vs-jev" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 30: ModernBERT versus Jev.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;But our ModernBERT model can’t play Tetris, for example, or do anything else besides movie classification without us fine-tuning it for the new task. But we could then throw a GPT-5.6 Luna or GPT-6 Luna model at it (which is a bit slower and also more expensive).&lt;/p&gt;    &lt;p&gt;A classic rule of thumb was:&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;Use a cheap LLM (like GPT-6 Luna) for one-off decision-making tasks;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;Fine-tune a custom classifier if we want to do this task repeatedly.&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;Now, something like Jev could take the place of the above to a) reduce latency and save money over Luna and b) save us the work of fine-tuning a custom model. (Although if we have a very high-volume task and we want to maximize speed and accuracy on a very specific task, it, of course, still makes sense to fine-tune.)&lt;/p&gt;    &lt;h2&gt;4. BERT- and GPT-style models with Jev API      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;    &lt;p&gt;Of course, we could also be adding a Jev-like API on top of a (Modern)BERT or any GPT-style model. Adding a Jev-like API is pretty straightforward. In fact, as soon as Jev was released, I built a Jev-like ModernBERT model to show how simple it is to build your own Jev model.&lt;/p&gt;    &lt;p&gt;I decided not to release my Jev clone because I changed my mind in the meantime. I.e., it’s trivial to put a Jev-like API on top of ModernBERT, and it’s trivial to fine-tune it on a bunch of classification tasks. But it’s not trivial to make this model work well on all different kinds of tasks (like Tetris) without extensive training and testing. (Also, there are already enough quick Jev clones riding on the hype train by now; the world doesn’t need another quick clone, but a strong open-weight version would be nice, of course.)&lt;/p&gt;    &lt;p&gt;However, if you are interested in how that retrofitting would work, here is a quick overview. For example, we can implement a Jev-like Choice API by adding a small classification head to any BERT-, GPT-, or T5-style model as illustrated earlier. But instead of having the output nodes in this head match the number of classes, we have it with only 1 output node, as illustrated for the modified GPT model below.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!Doyz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd09a503-5698-4cc1-8692-7a9bcb81ffba_1943x1931.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="gpt-1node" height="504.8598901098901" src="https://substackcdn.com/image/fetch/$s_!Doyz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd09a503-5698-4cc1-8692-7a9bcb81ffba_1943x1931.png" title="gpt-1node" width="508"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 31: GPT model where we replace the output layer (which maps to the whole vocabulary) with an output head with only one node.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Earlier, we discussed fine-tuning these models for a specific task like movie review classification with a pre-defined number of class labels (here, “positive” and “negative”). With this “1 output node” setup, we can actually extend this to a flexible and arbitrary number of classes.&lt;/p&gt;    &lt;p&gt;To extend this to an arbitrary number of classes (let’s consider the 3-class case of categorizing a customer ticket into the three categories “billing”, “technical”, “account”), we can do so as shown in the figure below.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!kK4c!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0ef4b0-2afd-4ba9-9be5-95440b50b2aa_6723x3497.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="jev-diy" height="757" src="https://substackcdn.com/image/fetch/$s_!kK4c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a0ef4b0-2afd-4ba9-9be5-95440b50b2aa_6723x3497.png" title="jev-diy" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 32: Using a BERT- or GPT-style model with a Jev-like choice API.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;As shown in the figure above, for each candidate option, we feed the model the input text, the task instructions, and the candidate’s description. The classification (scoring) head maps the output representation to a single scalar score. We then apply softmax across the candidate scores to obtain a probability distribution and return the highest-probability option.&lt;/p&gt;    &lt;p&gt;The main point here is that the proposed head has one output per class. Each class description produces its own representation, which the same head scores via the output layer (which is essentially a logistic regression model):&lt;/p&gt;    &lt;p&gt;where&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;          &lt;em&gt;s            &lt;sub&gt;i&lt;/sub&gt;&lt;/em&gt;is the scalar score for class label          &lt;em&gt;i&lt;/em&gt;;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;          &lt;em&gt;w&lt;/em&gt;is a learnable weight;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;          &lt;em&gt;b&lt;/em&gt;is a learnable bias unit;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;          &lt;em&gt;h            &lt;sub&gt;i&lt;/sub&gt;&lt;/em&gt;is the output of the transformer before the classification head.&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;For      &lt;em&gt;N&lt;/em&gt;candidates, this produces      &lt;em&gt;N&lt;/em&gt;scores. Softmax then returns      &lt;em&gt;N&lt;/em&gt;probabilities. The learned parameters and stay the same when      &lt;em&gt;N&lt;/em&gt;changes.&lt;/p&gt;    &lt;p&gt;Note that for      &lt;em&gt;h        &lt;sub&gt;i&lt;/sub&gt;&lt;/em&gt;:&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;BERT-style encoder models use the final hidden state of the          &lt;code&gt;[CLS]&lt;/code&gt;token (optionally passed through a pooling layer);&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;GPT-style autoregressive models typically use the final non-padding token’s hidden state, which we need because of the autoregressive nature as discussed earlier (the last non-padding token can attend to the preceding text, instructions, and candidate description).&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;Then, we fine-tune the model and the scoring head jointly using cross-entropy loss against the correct option. Since all candidates share the same scoring head, we can change the number and descriptions of the options without changing the architecture. But again, how well this works on unfamiliar tasks depends on the training data.&lt;/p&gt;    &lt;p&gt;By the way, why BERT- or GPT-style transformer-based models over simpler RNN and CNN architectures mentioned earlier? From a technical perspective, the approach outlined above works with either. But transformer-based models scale really well, meaning they can be pre-trained on larger datasets and benefit more than other models. Also, they can make good use of the information provided in the context (thanks to attention). So, if we want our model to generalize well to different target tasks without explicit fine-tuning on each one, pre-training a transformer-based model on a high-quality dataset likely gets us closer than RNN- or CNN-based models.&lt;/p&gt;    &lt;h2&gt;5. Jev architecture and training algorithm      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;    &lt;p&gt;So, how come Jev does so well on so many different tasks (from movie reviews to something arbitrary like playing Tetris)? Unfortunately, the architecture, algorithm, and training data details are not public. Also, keep in mind there’s a whole team and multi-million-dollar company behind it that specialized in and worked really hard on this model; we can’t expect to match that level of performance by training a ModernBERT-like model for a week on some open datasets.&lt;/p&gt;    &lt;p&gt;That being said, if I had to make an educated guess, architecture-wise, I’d guess that they are using something small similar to ModernBERT, hence, the low latency.&lt;/p&gt;    &lt;p&gt;For the training data, as mentioned before, the TypeSafe AI CEO      &lt;a href="https://x.com/CompleteSkeptic/status/2100617775823966680"&gt;said&lt;/a&gt;the following:&lt;/p&gt;    &lt;blockquote&gt;      &lt;p&gt;100% of our data is synthetic (but not the type of crap that is just spit out from an LLM obviously)&lt;/p&gt;&lt;/blockquote&gt;    &lt;p&gt;So, yeah, I think most of the effort went into curating this dataset. It’s something I’ve also been preaching to students and collaborators for many years. As a short anecdote, about 8 years ago, I was collaborating with another professor in the social sciences department and helped to design the experimental setup for a text classification problem. Her student spent many days hyperparameter tuning both a bag-of-words baseline and a BERT model to eke out ~2-5% accuracy. Then I suggested we each sit down for a few days to hand-label more data (I think our original dataset was around 300 samples, and we doubled the size), which resulted in a &amp;gt;10-20% accuracy boost. Yes, it’s important to      &lt;a href="https://rasbt.github.io/mlxtend/user_guide/plotting/plot_learning_curves/"&gt;plot learning curves&lt;/a&gt;for that purpose :).&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!RRQL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfa6e6db-1008-4418-b319-41f1bba35410_6020x2962.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="learning-curve" height="716" src="https://substackcdn.com/image/fetch/$s_!RRQL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfa6e6db-1008-4418-b319-41f1bba35410_6020x2962.png" title="learning-curve" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 33: Rule of thumb showing that often more data helps more than additional hyperparameter tuning.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Finally, the training algorithm. Again, the details are not known, but TypeSafe AI’s blog post      &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;states&lt;/a&gt;that they were using a new algorithm called Reinforcement Learning for Calibrated Decisions (RLCD):&lt;/p&gt;    &lt;blockquote&gt;      &lt;p&gt;“We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).”&lt;/p&gt;&lt;/blockquote&gt;    &lt;p&gt;This method is not public. A published method with a similar calibration objective is RLCR (Reinforcement Learning with Calibration Rewards), from the 2025 paper      &lt;a href="https://arxiv.org/abs/2507.16806"&gt;Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty&lt;/a&gt;. Before providing more details, the next section gives a brief overview of calibration in general.&lt;/p&gt;    &lt;h3&gt;5.1 On calibration      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Calibration is not a new topic and a step I recommend for any production model where you want to use and evaluate the class-membership probabilities. In short, calibration adjusts the model’s probability estimates so they better match observed class frequencies.&lt;/p&gt;    &lt;p&gt;For example, consider once more our IMDb movie classification problem. A model might assign a review 74% positive and 26% negative. These are probability estimates, but the model may be overconfident or underconfident. We cannot assume that the numerical values are reliable without evaluating calibration. I.e., predictions of 74% positive and 54% positive both yield the class label “positive” at a 50% threshold, although the model expresses more confidence in the 74% case.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!k414!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F642f91a5-0550-4d6d-8f70-85a1b77330c7_5046x2578.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="calibration" height="744" src="https://substackcdn.com/image/fetch/$s_!k414!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F642f91a5-0550-4d6d-8f70-85a1b77330c7_5046x2578.png" title="calibration" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 34: Illustration of how calibration works.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Calibration techniques use a separate labeled dataset, e.g., a held-out validation set, to adjust these probability estimates. After successful calibration, among many reviews assigned a positive probability of approximately 74%, roughly 74% should actually be positive. For more information, I recommend the good old      &lt;a href="https://scikit-learn.org/stable/modules/calibration.html"&gt;scikit-learn documentation&lt;/a&gt;.&lt;/p&gt;    &lt;p&gt;There are many approaches for this. For example, one that I used in my      &lt;a href="https://github.com/rasbt/ai-detector-from-scratch/blob/main/scripts/09_modernbert/modernbert.ipynb"&gt;AI detector project&lt;/a&gt;is temperature scaling. Temperature scaling divides the model’s logits by a temperature (      &lt;em&gt;T&lt;/em&gt;) value before applying softmax. We learn      &lt;em&gt;T&lt;/em&gt;(a positive number) by minimizing the cross-entropy loss on the calibration dataset while keeping the model weights fixed. A temperature above 1 reduces confidence in the highest-probability class, while a temperature between 0 and 1 increases it. We then use this same temperature for new predictions. Because this scaling preserves the ordering of the logits, the predicted class remains unchanged when we select the highest-probability class.&lt;/p&gt;    &lt;h3&gt;5.2 Reinforcement learning with calibration rewards (RLCR)      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Jev’s RLCD training method remains proprietary, but a related idea we can discuss is      &lt;em&gt;Reinforcement Learning with Calibration Rewards&lt;/em&gt;(RLCR), which was introduced in the 2025      &lt;a href="https://arxiv.org/abs/2507.16806"&gt;Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty&lt;/a&gt;paper. (Note that there is no officially established connection between the two methods, but I am assuming that they could be related.)&lt;/p&gt;    &lt;p&gt;Reinforcement learning with verifiable rewards (RLVR) typically rewards a correct answer with 1 and an incorrect answer with 0. (For a comprehensive explanation and implementation, I recommend checking out my      &lt;a href="https://amzn.to/4aAKiFY"&gt;Build A Reasoning Model (From Scratch) book&lt;/a&gt;:)).&lt;/p&gt;    &lt;p&gt;In short, RLCR adds an additional penalty for inaccurate confidence estimates.&lt;/p&gt;    &lt;p&gt;So, in the conventional RLVR method, the reward      &lt;em&gt;R&lt;/em&gt;is either 0 or 1 based on the answer correctness (we are ignoring an optional formatting reward and length penalty here, for simplicity).&lt;/p&gt;    &lt;p&gt;In RLCR, the model generates reasoning and an answer, which is followed by an uncertainty analysis and a numerical confidence (      &lt;em&gt;q&lt;/em&gt;). This modified reward is&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;\(R = c - (q - c)^2,\)&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;where is 1 for a correct answer and 0 otherwise. For example, an incorrect answer with 90% (0.9) confidence receives a reward of -0.81, since&lt;/p&gt;    &lt;div&gt;      &lt;div&gt;\(0 - (0.9 - 0)^2 = -0.81.\)&lt;/div&gt;&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;At 20% confidence, the reward is -0.04. And a correct answer at 90% confidence earns 0.99. (Readers familiar with evaluating calibrated models may notice that the squared-error term is the Brier penalty for the stated probability that the answer is correct.)&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!CodI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23d71def-2f99-4f25-a66e-9ee6e656c6d8_6070x2504.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="RLCR training loop" height="274.0824175824176" src="https://substackcdn.com/image/fetch/$s_!CodI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23d71def-2f99-4f25-a66e-9ee6e656c6d8_6070x2504.png" title="RLCR training loop" width="664"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 35: RLCR overview.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Note that the same (here, Qwen2.5-7B) model generates the uncertainty analysis and q as part of its response.&lt;/p&gt;    &lt;p&gt;The uncertainty analysis is a written assessment of where its answer could be wrong. After answering, the model examines missing evidence, ambiguous wording, questionable assumptions, or possible reasoning errors. The paper’s prompt asks it to identify specific uncertainties rather than propose corrections.&lt;/p&gt;    &lt;p&gt;Also, the system prompt explicitly asks for      &lt;em&gt;q&lt;/em&gt;, a number between 0 and 1 inside tags. It represents the model’s estimated probability that its answer is correct. The system generates all these parts sequentially in the same response.&lt;/p&gt;    &lt;p&gt;For illustration purposes, the model answer could be as follows:&lt;/p&gt;    &lt;blockquote&gt;      &lt;pre&gt;        &lt;code&gt;          &lt;code&gt;&amp;lt;think&amp;gt;...reasoning about the question...&amp;lt;/think&amp;gt;
&amp;lt;answer&amp;gt;positive&amp;lt;/answer&amp;gt;
&amp;lt;analysis&amp;gt;
The passage mentions a positive movie review, but the connection to the second paragraph is unclear.
&amp;lt;/analysis&amp;gt;
&amp;lt;confidence&amp;gt;0.6&amp;lt;/confidence&amp;gt;&lt;/code&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/blockquote&gt;    &lt;p&gt;Below is an annotated figure from the paper that illustrates how it improves the expected calibration error over regular RLVR with temperature scaling.&lt;/p&gt;    &lt;p&gt;On HotpotQA, RLCR reduces expected calibration error (ECE) from 0.37 to 0.03 compared with RLVR, with similar accuracy (62.1% versus 63.0%). Across six other datasets, average ECE falls from 0.46 to 0.21, while accuracy rises from 53.9% to 56.2%. These are the results in the paper’s Table 1(a).&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!0025!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0028832-3d38-4d48-93f4-fc4547db9fd0_6661x3645.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="RLCR accuracy and calibration results" height="797" src="https://substackcdn.com/image/fetch/$s_!0025!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0028832-3d38-4d48-93f4-fc4547db9fd0_6661x3645.png" title="RLCR accuracy and calibration results" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 36: RLCR results from the      &lt;a href="https://arxiv.org/abs/2507.16806"&gt;Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty&lt;/a&gt;paper.&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Conceptually, the “Classifier” in “RLVR + Classifier” is a supervised binary classifier that predicts whether the generated answer is correct.&lt;/p&gt;    &lt;p&gt;For Jev, I could apply calibration rewards directly to typed decisions, without generating reasoning text. For example, we could equip a GPT-style model with the classification head illustrated in Section 2.2 and adapt the RL training objective to reward accurate decisions and calibrated probabilities. We would train the backbone and classification head together. This would be an adaptation inspired by RLCR, and it remains unclear whether Jev’s RLCD works this way.&lt;/p&gt;    &lt;p&gt;Alternatively, since Direct Policy Optimization (DPO) is a cross-entropy (CE) alternative to Reinforcement Learning with Human Feedback (RLHF), we could also directly minimize CE + Brier loss on the classifier’s probabilities. I did this with the ModernBERT model, and there was a modest improvement over regular temperature scaling.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!LD_H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420c9874-1fe8-4a87-9c7d-483966e77879_6239x2423.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="rlcr-brier" height="565" src="https://substackcdn.com/image/fetch/$s_!LD_H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F420c9874-1fe8-4a87-9c7d-483966e77879_6239x2423.png" title="rlcr-brier" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 37: ModernBERT improvements via RLCR-inspired Brier-loss optimization. The ↑ symbol means higher is better. The ↓ symbol means lower is better.&lt;/div&gt;    &lt;h3&gt;5.3 Side note: Is calibration necessary?      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Cross-entropy already encourages accurate probability estimates. Under ideal conditions, that is enough to obtain calibrated predictions without an additional penalty.&lt;/p&gt;    &lt;p&gt;I.e., using a simple coin flip example analogous to the movie review classification, the two classes are “heads” and “tails.” If it’s a biased coin that lands heads 80% of the time (in expectation), the true probabilities are 80% heads and 20% tails. On a representative training set, when we optimize cross-entropy, the model will already learn these as a side-effect of minimizing cross-entropy (negative-log-likelihood).&lt;/p&gt;    &lt;p&gt;So, in short, if the model learns the true class probabilities, its predictions are calibrated and an additional penalty is unnecessary. (And Brier loss has the same theoretical optimum as cross-entropy, as described in      &lt;a href="https://www.tandfonline.com/doi/abs/10.1198/016214506000001437"&gt;Strictly Proper Scoring Rules, Prediction, and Estimation&lt;/a&gt;by Gneiting and Raftery, 2007.)&lt;/p&gt;    &lt;p&gt;In practice, however, we never train to the real optimum (in expectation) but train on a finite dataset.&lt;/p&gt;    &lt;p&gt;So, once a neural network correctly classifies most training examples, it can further reduce cross-entropy on the training set by becoming more confident. These confidence estimates may then generalize poorly to the test set and new data.&lt;/p&gt;    &lt;p&gt;For example, in      &lt;a href="https://proceedings.mlr.press/v70/guo17a.html"&gt;On Calibration of Modern Neural Networks&lt;/a&gt;, Guo et al. (2017) observed “neural networks can overfit to NLL without overfitting to the 0/1 loss.”&lt;/p&gt;    &lt;p&gt;That means test accuracy can still improve while test cross-entropy worsens. So overfitting can show up in the probability estimates even when classification accuracy continues to improve.&lt;/p&gt;    &lt;p&gt;Adding Brier loss changes how prediction errors are weighted during training. But whether this improves calibration needs to be checked on held-out data, of course. (The modest improvement in my experiment is an empirical observation for this setup, and we should not assume that adding Brier will always help.)&lt;/p&gt;    &lt;p&gt;In my experiments (see previous figure), adding the Brier loss to cross-entropy provided only a very small additional calibration benefit.&lt;/p&gt;    &lt;p&gt;But the motivation is more direct in RL contexts, such as described in RLCR, where we measure answer correctness and don’t already minimize cross-entropy. So, in RL contexts, the Brier term adds an incentive to report accurate confidence alongside the reward for answer correctness.&lt;/p&gt;    &lt;p&gt;&lt;/p&gt;    &lt;h2&gt;6. Who is Jev for?      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;    &lt;p&gt;Now that we covered Jev in reasonable detail, who is Jev for? Based on my assessment, it’s a pretty big target audience that intersects between people who a) want to save time by not having to fine-tune a custom classifier for every decision task and b) want to save money over using “the big guns” like GPT-6 and similar LLMs.&lt;/p&gt;    &lt;p&gt;There are also lots of interesting use cases to think of. In the section below is a short list of ideas.&lt;/p&gt;    &lt;h3&gt;6.1 Jev use cases      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;The list of use cases is seemingly endless. Sure, there are flashy examples like having it play video games like Tetris (as I showed in my earlier video above to demonstrate the low latency of the API). However, sorting emails is a more practical choice for most. For example, one could use Jev as an additional spam filter, prioritization, and so on.&lt;/p&gt;    &lt;p&gt;But beyond things like that, it could also serve as a tool to augment LLM-agent harnesses. Here it could be used as&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;        &lt;p&gt;a pre-screener for regular LLM-agent harnesses to scan our contexts for prompt injection;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;select the reasoning effort level for a model;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;select a skill.md from your registered skill library;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;use it as a judge for evaluation or self-refinement;&lt;/p&gt;&lt;/li&gt;      &lt;li&gt;        &lt;p&gt;find relevant files for context building.&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;The list is practically endless. Someone even posted a      &lt;a href="https://arxiv.org/abs/2609.30216"&gt;survey on arXiv&lt;/a&gt;analyzing 2,170 Jev-related projects that have sprung up recently in just a few days.&lt;/p&gt;    &lt;h3&gt;6.2 About Jev clones and local Jevs      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h3&gt;    &lt;p&gt;Since Jev’s launch, 100s, if not 1000s, of quick Jev clones have sprung up. Most of the ones I checked out were just quick examples fine-tuning a ModernBERT or Qwen model with SFT and adding a Jev-like API.&lt;/p&gt;    &lt;p&gt;Based on what I can tell, none of them achieves the same level of performance as Jev on such a breadth of tasks. Comparing these projects to Jev is like comparing Alpaca (the early instruction-finetuned LLM based off of the original LLaMA open-weight model) to GPT-6. Sure, it may work similarly in spirit, but your mileage will vary on real-world tasks.&lt;/p&gt;    &lt;p&gt;No offense, but they seem like quick projects to jump on the hype train, and I don’t want to plug a specific one here.&lt;/p&gt;    &lt;p&gt;One exception here might be the      &lt;a href="https://github.com/urchade/GLiNER"&gt;GLiNER&lt;/a&gt;project, which has been around for ~3 years ago, and while it’s not fundamentally the same, it can be used for similar things. However, based on a      &lt;a href="https://github.com/AbdelStark/jev-benchmarks/blob/main/results/reports/btzsc-pilot-v1.md"&gt;benchmark&lt;/a&gt;, Jev is definitely stronger.&lt;/p&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!1X1i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6009d1b8-e300-41b8-bf62-3eb10532025f_4065x1743.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="gliner" height="624" src="https://substackcdn.com/image/fetch/$s_!1X1i!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6009d1b8-e300-41b8-bf62-3eb10532025f_4065x1743.png" title="gliner" width="1456"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 38: GLiNER vs Jev comparison via      &lt;a href="https://github.com/AbdelStark/jev-benchmarks/blob/main/results/reports/btzsc-pilot-v1.md"&gt;https://github.com/AbdelStark/jev-benchmarks/blob/main/results/reports/btzsc-pilot-v1.md&lt;/a&gt;&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;Anyways, there is definitely a genuine need and desire for a good open-weight alternative to Jev. The reason is not necessarily cost (because Jev is so cheap) but privacy. (Also, if Jev is already that fast over the internet, imagine how fast it could be on beefy local hardware.)&lt;/p&gt;    &lt;p&gt;So, of course I’d like a DeepSeek (or Kimi, or GLM, or MiMo) moment for Jev. I am sure people are working on this, but it will take some time. TypeSafe AI worked on this for many months with a dedicated team of experts. It’s unrealistic to expect to replicate that development and evaluation work in a week. (However, since this is a smaller and more efficient model than a 1-trillion-parameter model, by nature, it hopefully won’t take that long.)&lt;/p&gt;    &lt;p&gt;      &lt;strong&gt;Update 1 (29 September, 10:20 am PT):&lt;/strong&gt;OpenAI      &lt;a href="https://openai.com/index/devday-2026-recap/"&gt;just announced&lt;/a&gt;the Decision API at their DevDay 2026 conference (29 September). It appears to be a Jev-like model directly integrated into their platform.&lt;/p&gt;    &lt;blockquote&gt;      &lt;p&gt;Decisions API enables real-time decision-making by focusing Luna’s intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent’s next action.&lt;/p&gt;      &lt;p&gt;Available in limited preview today with a broad release planned in the coming days.&lt;/p&gt;&lt;/blockquote&gt;    &lt;div&gt;      &lt;a href="https://substackcdn.com/image/fetch/$s_!nILp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecf4d0d-2865-4ad1-8322-24f27fc3d323_1840x876.png" target="_blank"&gt;        &lt;div&gt;          &lt;img alt="" height="262.7307692307692" src="https://substackcdn.com/image/fetch/$s_!nILp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffecf4d0d-2865-4ad1-8322-24f27fc3d323_1840x876.png" width="552"&gt;&lt;/img&gt;          &lt;div&gt;            &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/a&gt;Figure 39: OpenAI’s Jev-like Decision API announced at the OpenAI Dev Day.&lt;/div&gt;    &lt;p&gt;      &lt;strong&gt;Update 2:&lt;/strong&gt;A reader shared the      &lt;a href="https://contrastive-lm.notion.site/"&gt;Contrastive Language Models&lt;/a&gt;project with me in the comments below. It’s one of the nicer-looking Jev-likes. When I tried it, its IMDb test accuracy is just 82.90% (over 96.47% in Jev), and it      &lt;a href="https://sebastianraschka.com/videos/tetris-clm"&gt;fails the Tetris test&lt;/a&gt;.&lt;/p&gt;    &lt;p&gt;Another reader wrote in about      &lt;a href="https://huggingface.co/convaiinnovations/laya"&gt;Laya&lt;/a&gt;, another strong contender. It did better at IMDb (92.33% test accuracy) but then      &lt;a href="https://sebastianraschka.com/videos/tetris-laya/"&gt;failed the Tetris test&lt;/a&gt;even worse.&lt;/p&gt;    &lt;p&gt;So, my point still stands. There are many solid alternatives. But what makes Jev so popular is not a unique idea (as mentioned before, classifiers existed for a long time) but the fact that it performs so well at so many diverse tasks. The only task where it allegedly doesn’t do well (as a reader told me) is legal reviews in non-English languages. (But the weak multilingual performance can be circumvented by connecting it to a cheap translation model, e.g., GPT -6 Luna with a small latency overhead.)&lt;/p&gt;    &lt;h2&gt;Conclusion      &lt;div&gt;        &lt;div&gt;          &lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/h2&gt;    &lt;p&gt;All in all, at first glance, Jev doesn’t seem to offer anything fundamentally new. After all, one might say that “it’s just a classifier” (with a nice API on top of it). However, it works surprisingly well across a huge range of tasks. In that sense, it’s the plug-and-play version of dedicated classifiers.&lt;/p&gt;    &lt;p&gt;Sure, experts will still fine-tune specialist models, but the bar to justify fine-tuning a specialized model is now much higher, since it’s easy to just throw a Jev or Jev-like at it and get good-enough results.&lt;/p&gt;    &lt;p&gt;Also, I don’t expect Jev and Jev-likes to unlock new capabilities or solve previously unsolvable tasks. However, if we make such models part of our agent harnesses to aid the expensive GPT-6 or Opus 5.5 models in decision-making within that harness, using agent harnesses could become much faster and cheaper in the future.&lt;/p&gt;    &lt;p&gt;&lt;/p&gt;    &lt;div&gt;      &lt;hr&gt;&lt;/hr&gt;&lt;/div&gt;    &lt;p&gt;If you found this article useful, consider becoming a paid      &lt;a href="https://magazine.sebastianraschka.com/subscribe"&gt;subscriber&lt;/a&gt;to Ahead of AI. Your support helps me spend more time on the experiments and detailed illustrations behind articles like this.&lt;/p&gt;    &lt;p&gt;You can also support my work through my books,      &lt;a href="https://amzn.to/4fqvn0D"&gt;Build a Large Language Model (From Scratch)&lt;/a&gt;and      &lt;a href="https://amzn.to/4aAKiFY"&gt;Build a Reasoning Model (From Scratch)&lt;/a&gt;, if you’d like to implement these concepts yourself.&lt;/p&gt;    &lt;p&gt;      &lt;strong&gt;Thanks for reading and supporting my independent research!&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63290-%E6%96%87%E6%9C%AC-%E5%88%86%E7%B1%BB-%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B</guid>
      <pubDate>Wed, 30 Sep 2026 09:26:32 CST</pubDate>
    </item>
    <item>
      <title>别吹 Demo 了：读完 Pi 的 Harness v2 和 PR</title>
      <link>https://itindex.net/detail/63289-demo-pi-harness</link>
      <description>&lt;h1&gt;别吹 Demo 了：读完 Pi 的 Harness v2 和 PR #8172，聊聊写 Agent 的那些烂坑&lt;/h1&gt;
 &lt;p&gt;现在的 Agent 圈子很浮躁。&lt;/p&gt;
 &lt;p&gt;大家喜欢看炫酷的控制台动效，喜欢接几十个工具，喜欢等下一个更强的大模型。&lt;/p&gt;
 &lt;p&gt;但只要你把 Agent 放进真实的工程环境跑几天，就会发现现实很残酷：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;跑了十分钟的任务，网络一抖或者内存一爆，进程挂了，会话直接报废。&lt;/li&gt;
  &lt;li&gt;执行一个    &lt;code&gt;npm test&lt;/code&gt;，吐出三万行日志，上下文当场被撑爆，模型开始胡言乱语。&lt;/li&gt;
  &lt;li&gt;为了让用户实时插话，框架往会话中间塞了一条消息，底层的 KV Cache 全废了，Token 费用翻了五倍，响应慢得像蜗牛。&lt;/li&gt;
  &lt;li&gt;搞多 Agent 并发，多个任务抢着写同一个状态文件，数据直接写坏。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;模型不是瓶颈，Demo 也没有意义。包裹在模型外面的 Harness（执行底盘）太烂，才是 Agent 落不了地的根本原因。&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;最近，  &lt;strong&gt;Pi（   &lt;code&gt;pi.dev&lt;/code&gt; /    &lt;code&gt;earendil-works/pi&lt;/code&gt;）&lt;/strong&gt; 发布了   &lt;code&gt;harness-v2.md&lt;/code&gt; 规范，随后合入了   &lt;strong&gt;PR #8172&lt;/strong&gt;。这是一份没有废话的工业级 Agent 底盘设计。&lt;/p&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h2&gt;1. Pi 是谁？&lt;/h2&gt;
 &lt;p&gt;在开源 Agent 项目里，Pi 很低调，但地位很硬。&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;   &lt;strong&gt;核心作者与团队&lt;/strong&gt;：Mario Zechner（GitHub:    &lt;code&gt;@badlogic&lt;/code&gt;，著名开源游戏引擎    &lt;strong&gt;libGDX&lt;/strong&gt; 的缔造者，典型的老派系统级程序员）与核心团队成员 David Brailovsky（   &lt;code&gt;@davidbrai&lt;/code&gt;）、Vegar Stikbakke（   &lt;code&gt;@vegarsti&lt;/code&gt;）。在设计文档中甚至明确写着：   &lt;em&gt;“如果设计不成立，停下来在 Discord 上找 Mario 商讨。”&lt;/em&gt;&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;PR #8172 作者&lt;/strong&gt;：Adam Teale（   &lt;code&gt;@adamteale&lt;/code&gt;）。这个仓库门槛极高，非协作者提 PR 会被 Bot 秒关，只有 Maintainer 明确给    &lt;code&gt;lgtm&lt;/code&gt; 才能进入流程。代码质量把控极严。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;项目定位&lt;/strong&gt;：   &lt;strong&gt;“Primitives, not features”（只做原语，不做死板特性）&lt;/strong&gt;。它不搞花哨的 UI，只提供    &lt;code&gt;@earendil-works/pi-agent-core&lt;/code&gt;（运行时）、   &lt;code&gt;pi-ai&lt;/code&gt;（模型层）和    &lt;code&gt;pi-tui&lt;/code&gt;（终端差分渲染）。像 OpenClaw 这类复杂的自主 Agent 框架，底层直接依赖 Pi 的 SDK。&lt;/li&gt;
&lt;/ul&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h2&gt;2. 竞品锐评：市面上的 Harness 错在哪？&lt;/h2&gt;
 &lt;p&gt;把 Pi 的设计和主流竞品放在一起看，很多流行框架的做法其实很业余。&lt;/p&gt;
 &lt;h3&gt;锐评 LangChain / LangGraph：自找麻烦的“图状态机”&lt;/h3&gt;
 &lt;p&gt;LangGraph 试图用有向图和 Generator 状态机去模拟每一步执行，搞出一堆复杂的类型转换和模板代码。一旦中途报错，状态机极难恢复。&lt;/p&gt;
 &lt;blockquote&gt;
  &lt;p&gt;   &lt;strong&gt;Pi 的选择&lt;/strong&gt;：Pi 早期也写过一套 Generator 状态机（   &lt;code&gt;harness-v2-generator.md&lt;/code&gt;），后来直接废弃。Mario 的理由很简单：   &lt;strong&gt;直线的     &lt;code&gt;async/await&lt;/code&gt; 最容易调试和理解。&lt;/strong&gt; 不需要把代码切成碎块，只要把副作用边界划清就行。&lt;/p&gt;
&lt;/blockquote&gt;
 &lt;h3&gt;锐评 DeepSeek Harness (dsh)：粗暴丢数据的截断器&lt;/h3&gt;
 &lt;p&gt;处理工具返回的超大文本时，DeepSeek 的 Harness（dsh）采用固定截断策略（保留头 4096 字符、尾 1024 字符，中间直接扔掉）。&lt;/p&gt;
 &lt;blockquote&gt;
  &lt;p&gt;   &lt;strong&gt;PR #8172 的爆破&lt;/strong&gt;：PR #8172 作者 Adam Teale 指出：   &lt;strong&gt;dsh 的做法是永久丢失数据。&lt;/strong&gt; 报错堆栈的关键信息通常就在中间，扔掉之后模型就彻底瞎了。PR #8172 提出了    &lt;strong&gt;Spill-copy-on-prune&lt;/strong&gt;：裁剪 Context 的同时，把完整内容写入磁盘文件，并保留精确的字符偏移。模型不仅知道中间被裁了，还能用    &lt;code&gt;grep&lt;/code&gt; 去磁盘文件里把那几行找出来。&lt;/p&gt;
&lt;/blockquote&gt;
 &lt;h3&gt;锐评 AutoGPT / CrewAI：没有容灾的纸糊玩具&lt;/h3&gt;
 &lt;p&gt;这些项目擅长展示“多智能体开会”，但根本没有真正的持久化与崩溃恢复（Durability）。任务跑了 30 步，第 29 步崩溃，整局重来。&lt;/p&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h2&gt;3. Agent 开发者天天踩的五个坑，Pi 怎么解？&lt;/h2&gt;
 &lt;h3&gt;坑 1：中途插消息，KV Cache 当场报废&lt;/h3&gt;
 &lt;p&gt;  &lt;strong&gt;现象&lt;/strong&gt;：
Agent 正在跑工具，用户发了一条修正指令（Steer）。很多框架直接把这条消息插到对话历史里。
大模型的 Prompt Cache 依赖严格的前缀匹配。你在中间动一个字，后面的缓存全部失效。费用暴涨，延迟翻倍。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;Pi 的解法：Append-Only Context（只增上下文）&lt;/strong&gt;
Pi 确立了一条铁律：  &lt;strong&gt;上下文只能在尾部增长。&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;运行中产生的配置变动、用户插话、延迟写入，先暂存到内存队列。&lt;/li&gt;
  &lt;li&gt;等当前回合（Turn）跑完，到达    &lt;strong&gt;Checkpoint（检查点）&lt;/strong&gt; 时，统一追加到上下文末尾。&lt;/li&gt;
  &lt;li&gt;缓存前缀永远不变，KV Cache 命中率永远最高。&lt;/li&gt;
&lt;/ul&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h3&gt;坑 2：进程挂了，留下半截死状态&lt;/h3&gt;
 &lt;p&gt;  &lt;strong&gt;现象&lt;/strong&gt;：
模型发起了两个工具调用，第一个执行完了，进程崩溃。重启后，对话历史里有调用指令却没有工具返回，下一次请求直接报 API 格式错误。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;Pi 的解法：WAL 意图日志与预分配 ID&lt;/strong&gt;
Pi 借用了数据库系统的 Write-Ahead Logging 思想：&lt;/p&gt;
 &lt;ol&gt;
  &lt;li&gt;   &lt;strong&gt;执行工具前&lt;/strong&gt;：先写一条    &lt;code&gt;step_attempt&lt;/code&gt; 意图记录到磁盘，里面包含预先分配好的结果 ID（   &lt;code&gt;resultEntryId&lt;/code&gt;）。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;执行完成后&lt;/strong&gt;：用这个 ID 写入真实的工具结果。&lt;/li&gt;
&lt;/ol&gt;
 &lt;pre&gt;  &lt;code&gt;interface StepAttemptRecord extends RecordBase {
  type: &amp;quot;step_attempt&amp;quot;;
  runId: string;
  step: &amp;quot;assistant&amp;quot; | &amp;quot;compaction&amp;quot; | &amp;quot;branch_summary&amp;quot;;
  attempt: number;        // 重试次数写在磁盘上，进程重启也不会无限重试
  resultEntryId: string; // 预先指定的 Entry ID
}
&lt;/code&gt;&lt;/pre&gt;
 &lt;p&gt;系统重启时，恢复程序扫描日志。如果发现有意图记录但没有对应的结果，就能精准知道崩溃发生在哪个位置：能重试的自动重试，不能重试的自动写入一条   &lt;code&gt;&amp;quot;interrupted&amp;quot;&lt;/code&gt; 占位，  &lt;strong&gt;绝不会出现残缺的孤儿状态。&lt;/strong&gt;&lt;/p&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h3&gt;坑 3：超大输出撑爆上下文，盲目截断导致模型幻觉&lt;/h3&gt;
 &lt;p&gt;  &lt;strong&gt;现象&lt;/strong&gt;：
执行一个命令吐出 50KB 日志，塞进上下文直接超限；如果直接把中间砍掉，模型看不到错误信息，开始胡乱猜测。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;Pi 的解法（PR #8172 三级无损管道）&lt;/strong&gt;：&lt;/p&gt;
 &lt;ol&gt;
  &lt;li&gt;   &lt;strong&gt;&amp;gt; 50K 字符&lt;/strong&gt;：内容全部写入磁盘文件（   &lt;code&gt;~/.pi/agent/cache/tool-spill/&lt;/code&gt;），上下文里只给简短预览和文件路径。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;&amp;gt; 8K 字符&lt;/strong&gt;：
   &lt;ul&gt;
    &lt;li&gt;完整内容写入磁盘。&lt;/li&gt;
    &lt;li&gt;上下文保留头 4096 字符 + 尾 1024 字符。&lt;/li&gt;
    &lt;li&gt;     &lt;strong&gt;在字符串索引 0 处&lt;/strong&gt;写入标记      &lt;code&gt;[pruned — full at /path...]&lt;/code&gt;，并标明被裁剪的精确范围      &lt;code&gt;[start, end)&lt;/code&gt;。&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;防死循环&lt;/strong&gt;：模型调用读取工具去读溢出文件时，Pruner 自动放行，绝不二次截断。&lt;/li&gt;
&lt;/ol&gt;
 &lt;p&gt;  &lt;strong&gt;实战测试数据&lt;/strong&gt;（基于 GLM-5.3 和 DeepSeek-V4-Flash 的 19 次真实测试）：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;未命中缓存的 Prefill Token 减少    &lt;strong&gt;72% ~ 88%&lt;/strong&gt;。&lt;/li&gt;
  &lt;li&gt;每轮请求上下文占用减少    &lt;strong&gt;26% ~ 35%&lt;/strong&gt;。&lt;/li&gt;
  &lt;li&gt;幻觉测试通过率 100%（0/9 幻觉），模型能通过精确偏移用    &lt;code&gt;grep&lt;/code&gt;/   &lt;code&gt;sed&lt;/code&gt; 从磁盘精准读回丢失的日志。&lt;/li&gt;
&lt;/ul&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h3&gt;坑 4：多任务并发，状态文件被写烂&lt;/h3&gt;
 &lt;p&gt;  &lt;strong&gt;现象&lt;/strong&gt;：
多 Agent 系统让多个子任务并发修改同一个状态文件，在 Node.js 或 Python 异步环境下极易产生写覆盖和文件损坏。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;Pi 的解法：Lanes（泳道）与不可变树&lt;/strong&gt;
Pi 把会话结构分成了两层：&lt;/p&gt;
 &lt;pre&gt;  &lt;code&gt; 会话树（只增、共享、无状态）:   a ─── b ─── c ─── d
                                       └─── e ─── f
 
 泳道（各自独立）:               main ──&amp;gt; 指向 d
                                slack_thread_1 ──&amp;gt; 指向 f
&lt;/code&gt;&lt;/pre&gt;
 &lt;ul&gt;
  &lt;li&gt;   &lt;strong&gt;底层树（Tree）&lt;/strong&gt;：所有节点只增不减，所有泳道共享，只读不改。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;泳道（Lane）&lt;/strong&gt;：相当于 Git 分支。每个 Lane 只有自己的指针和操作队列，两个 Lane 并行工作完全不需要加锁互斥。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;单写者（Single Writer）&lt;/strong&gt;：整个会话只有一个写入器，所有并发写操作通过单调递增的序列号排队追加。&lt;/li&gt;
&lt;/ul&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h3&gt;坑 5：异步 Check-Then-Act 竞态&lt;/h3&gt;
 &lt;p&gt;  &lt;strong&gt;现象&lt;/strong&gt;：
代码判断“当前任务已结束”，准备关闭会话；就在这一瞬间，用户发来了   &lt;code&gt;abort&lt;/code&gt; 或新消息。时序交错，导致中止失败或消息丢失。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;Pi 的解法：Lane Mutation Line（泳道原子排队线）&lt;/strong&gt;
Pi 用一条极简的 Promise 链解决了所有竞态问题：&lt;/p&gt;
 &lt;pre&gt;  &lt;code&gt;let tail: Promise&amp;lt;unknown&amp;gt; = Promise.resolve();

function mutateLane&amp;lt;T&amp;gt;(job: () =&amp;gt; Promise&amp;lt;T&amp;gt;): Promise&amp;lt;T&amp;gt; {
  const result = tail.then(job);
  tail = result.then(() =&amp;gt; undefined, () =&amp;gt; undefined);
  return result;
}
&lt;/code&gt;&lt;/pre&gt;
 &lt;ul&gt;
  &lt;li&gt;状态检查和落盘写入，必须打包成一个同步 Job 放进队列执行。&lt;/li&gt;
  &lt;li&gt;网络请求、模型调用、工具执行在队列外面跑，跑完了再排队写状态。&lt;/li&gt;
  &lt;li&gt;任何两个并发操作，在时间线上永远只有    &lt;code&gt;[先 A 后 B]&lt;/code&gt; 或    &lt;code&gt;[先 B 后 A]&lt;/code&gt;，没有第三种可能。&lt;/li&gt;
&lt;/ul&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h2&gt;4. 确定性测试：线上代码怎么写，测试就怎么跑&lt;/h2&gt;
 &lt;p&gt;很多框架的单测都是在 Mock 数据，一到线上就出 Bug。&lt;/p&gt;
 &lt;p&gt;Pi 定义了一个纯粹的副作用接口   &lt;code&gt;Effects&lt;/code&gt;（  &lt;code&gt;fx&lt;/code&gt;）：&lt;/p&gt;
 &lt;pre&gt;  &lt;code&gt;interface Effects {
  appendEntry(...): Promise&amp;lt;Entry&amp;gt;;
  appendRecord(...): Promise&amp;lt;T&amp;gt;;
  streamAssistant(...): Promise&amp;lt;SettledAssistantMessage&amp;gt;;
  executeTool(...): Promise&amp;lt;{ result: AgentToolResult; isError: boolean }&amp;gt;;
}
&lt;/code&gt;&lt;/pre&gt;
 &lt;p&gt;Agent 的所有逻辑，只能通过   &lt;code&gt;fx&lt;/code&gt; 操作外部世界。&lt;/p&gt;
 &lt;p&gt;这带来了一个巨大的优势：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;   &lt;strong&gt;生产环境（    &lt;code&gt;drive: &amp;quot;automatic&amp;quot;&lt;/code&gt;）&lt;/strong&gt;：   &lt;code&gt;fx&lt;/code&gt; 直通系统，异步全速运行。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;测试环境（    &lt;code&gt;drive: &amp;quot;manual&amp;quot;&lt;/code&gt;）&lt;/strong&gt;：每一个操作都会在    &lt;code&gt;fx&lt;/code&gt; 边界停住。测试脚本可以像单步调试器一样，一步一步推着 Agent 走——可以在执行工具前强制断电，也可以在两个工具调用中间强行插入中断信号。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;测试测的，就是线上跑的原版代码。&lt;/strong&gt;&lt;/p&gt;
 &lt;hr&gt;&lt;/hr&gt;
 &lt;h2&gt;5. 总结&lt;/h2&gt;
 &lt;p&gt;做 Agent，写 Prompt 只是第一步。&lt;/p&gt;
 &lt;p&gt;当系统进入真实工程，决定生死的是这几件事：&lt;/p&gt;
 &lt;ol&gt;
  &lt;li&gt;   &lt;strong&gt;保住 KV Cache&lt;/strong&gt;：不要随意插队改前缀。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;做好 WAL 与崩溃恢复&lt;/strong&gt;：把每一次动作先记下来再执行。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;无损管理上下文&lt;/strong&gt;：大输出落盘，保留偏移，给模型找回数据的能力。&lt;/li&gt;
  &lt;li&gt;   &lt;strong&gt;管好并发与时序&lt;/strong&gt;：单写者、只增树、排队写。&lt;/li&gt;
&lt;/ol&gt;
 &lt;p&gt;把这些脏活累活做干净，Agent 才能从玩具变成工具。&lt;/p&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>post</category>
      <guid isPermaLink="true">https://itindex.net/detail/63289-demo-pi-harness</guid>
      <pubDate>Mon, 24 Aug 2026 04:45:00 CST</pubDate>
    </item>
    <item>
      <title>微软官方开源：一键把全新 Windows 11 变成干净、清爽的开发机</title>
      <link>https://itindex.net/detail/63288-%E5%BE%AE%E8%BD%AF-%E5%AE%98%E6%96%B9-%E5%BC%80%E6%BA%90</link>
      <description>&lt;p&gt;  &lt;strong&gt;重装 Windows、换新电脑，或者新建一台虚拟机&lt;/strong&gt;，在安装好系统之后，最麻烦的还是配置环境：关闭 Widgets 和各种推荐，打开文件扩展名，装 Git、VS Code、Python、Node.js，再配置 WSL、Ubuntu……&lt;/p&gt;



 &lt;p&gt;这些事情都不难，就是又多又碎，又容易遗漏。&lt;/p&gt;



 &lt;p&gt;前段时间，微软开源了一个神奇的工具：  &lt;strong&gt;Windows Developer Config&lt;/strong&gt;，按照微软自己的说法就是：&lt;/p&gt;



 &lt;p&gt;  &lt;strong&gt;一次操作即可将一台全新的 Windows 11 电脑变成一个干净、无干扰的开发工作站&lt;/strong&gt;&lt;/p&gt;



 &lt;img alt="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 1" src="https://www.appinn.com/wp-content/uploads/2026/09/Copy-of-appinn-homework-2026-09-27T200700.776.jpg" title="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 1"&gt;&lt;/img&gt;



 &lt;h2&gt;标准安装&lt;/h2&gt;



 &lt;p&gt;这款叫做   &lt;strong&gt;Windows Developer Config&lt;/strong&gt; 的开源工具，用起来简单粗暴，只需要在 PowerShell 中运行：&lt;/p&gt;


 &lt;div&gt;  &lt;pre&gt;
irm https://aka.ms/devconfig/standard/setup.ps1 | iex
&lt;/pre&gt;&lt;/div&gt;


 &lt;p&gt;就可以实现：&lt;/p&gt;



 &lt;h3&gt;  &lt;strong&gt;删除/关闭&lt;/strong&gt;&lt;/h3&gt;



 &lt;ul&gt;
  &lt;li&gt;减少开始菜单、Windows Search 中的推荐和干扰内容&lt;/li&gt;



  &lt;li&gt;调整文件资源管理器中的最近项目、推荐等内容&lt;/li&gt;



  &lt;li&gt;减少部分 Windows 默认界面中的干扰项&lt;/li&gt;
&lt;/ul&gt;



 &lt;h3&gt;  &lt;strong&gt;安装/配置&lt;/strong&gt;&lt;/h3&gt;



 &lt;ul&gt;
  &lt;li&gt;   &lt;strong&gt;VS Code、Git、GitHub CLI、GitHub Copilot CLI&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;Python 3.14 + uv&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;Node.js LTS + nvm&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;.NET SDK 10&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;Windows Terminal、PowerShell 7、Oh My Posh&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;WSL + Ubuntu&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;   &lt;strong&gt;PowerToys、Coreutils for Windows、Windows App CLI&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;配置 Cascadia Mono NF 字体、深色模式、长路径等 Windows 设置&lt;/li&gt;
&lt;/ul&gt;



 &lt;p&gt;这是一套完整的开发环境，再也不需要手动设置了。&lt;/p&gt;



 &lt;img alt="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 2" src="https://www.appinn.com/wp-content/uploads/2026/09/ChatGPT-2026-9-27-20_08_59.avif" title="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 2"&gt;&lt;/img&gt;



 &lt;hr&gt;&lt;/hr&gt;



 &lt;h2&gt;完整安装&lt;/h2&gt;



 &lt;p&gt;而如果使用更激进的   &lt;strong&gt;Full&lt;/strong&gt; 配置：&lt;/p&gt;


 &lt;div&gt;  &lt;pre&gt;
irm https://aka.ms/devconfig/full/setup.ps1 | iex
&lt;/pre&gt;&lt;/div&gt;


 &lt;p&gt;就可以实现：&lt;/p&gt;



 &lt;h3&gt;  &lt;strong&gt;关闭/&lt;/strong&gt;调整&lt;/h3&gt;



 &lt;ul&gt;
  &lt;li&gt;关闭    &lt;strong&gt;Widgets 小组件&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭 Windows Search 中的    &lt;strong&gt;Web 搜索结果&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭    &lt;strong&gt;搜索亮点（Search Highlights）&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭开始菜单中的   &lt;strong&gt;推荐内容&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭开始菜单中的   &lt;strong&gt;账户通知&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭文件资源管理器「快速访问」中的   &lt;strong&gt;常用文件夹&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭「快速访问」中的   &lt;strong&gt;最近文件&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭文件资源管理器中的   &lt;strong&gt;推荐文件、云端文件&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭文件资源管理器中的   &lt;strong&gt;同步提供商提示&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;隐藏   &lt;strong&gt;蓝牙托盘图标&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;关闭    &lt;strong&gt;PowerToys Always On Top&lt;/strong&gt; 的通知&lt;/li&gt;



  &lt;li&gt;Edge 新标签页改成   &lt;strong&gt;空白页&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;跳过 Edge    &lt;strong&gt;首次运行欢迎界面&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;全局开启   &lt;strong&gt;勿扰模式&lt;/strong&gt;，关闭 Toast 通知&lt;/li&gt;
&lt;/ul&gt;



 &lt;h3&gt;  &lt;strong&gt;安装&lt;/strong&gt;&lt;/h3&gt;



 &lt;ul&gt;
  &lt;li&gt;Windows Terminal&lt;/li&gt;



  &lt;li&gt;Intelligent Terminal&lt;/li&gt;



  &lt;li&gt;PowerShell 7&lt;/li&gt;



  &lt;li&gt;Git&lt;/li&gt;



  &lt;li&gt;GitHub CLI&lt;/li&gt;



  &lt;li&gt;Azure CLI&lt;/li&gt;



  &lt;li&gt;GitHub Copilot CLI&lt;/li&gt;



  &lt;li&gt;Visual Studio Code&lt;/li&gt;



  &lt;li&gt;.NET SDK 10&lt;/li&gt;



  &lt;li&gt;Python 3.14&lt;/li&gt;



  &lt;li&gt;uv&lt;/li&gt;



  &lt;li&gt;Node.js LTS&lt;/li&gt;



  &lt;li&gt;nvm for Windows&lt;/li&gt;



  &lt;li&gt;Coreutils for Windows&lt;/li&gt;



  &lt;li&gt;Oh My Posh&lt;/li&gt;



  &lt;li&gt;Windows App CLI&lt;/li&gt;



  &lt;li&gt;PowerToys&lt;/li&gt;



  &lt;li&gt;WSL + Ubuntu&lt;/li&gt;
&lt;/ul&gt;



 &lt;p&gt;  &lt;strong&gt;同时配置：&lt;/strong&gt;&lt;/p&gt;



 &lt;ul&gt;
  &lt;li&gt;开启    &lt;strong&gt;Developer Mode&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;开启    &lt;strong&gt;Windows Sudo&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;开启 Win32    &lt;strong&gt;长路径支持&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;开启    &lt;strong&gt;Remote Desktop&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;显示文件扩展名&lt;/li&gt;



  &lt;li&gt;显示隐藏文件&lt;/li&gt;



  &lt;li&gt;文件资源管理器默认打开「此电脑」&lt;/li&gt;



  &lt;li&gt;标题栏显示完整路径&lt;/li&gt;



  &lt;li&gt;任务栏右键增加「结束任务」&lt;/li&gt;



  &lt;li&gt;Windows 和应用切换为   &lt;strong&gt;深色模式&lt;/strong&gt;&lt;/li&gt;



  &lt;li&gt;PowerShell 7 设为默认终端配置&lt;/li&gt;



  &lt;li&gt;配置 Oh My Posh&lt;/li&gt;



  &lt;li&gt;安装并使用 Cascadia Mono NF 字体&lt;/li&gt;



  &lt;li&gt;配置 GitHub Copilot Terminal Profile&lt;/li&gt;
&lt;/ul&gt;



 &lt;img alt="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 3" src="https://www.appinn.com/wp-content/uploads/2026/09/ChatGPT-2026-9-27-20_09_49.avif" title="&amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#24320;&amp;#28304;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#25226;&amp;#20840;&amp;#26032; Windows 11 &amp;#21464;&amp;#25104;&amp;#24178;&amp;#20928;&amp;#12289;&amp;#28165;&amp;#29245;&amp;#30340;&amp;#24320;&amp;#21457;&amp;#26426; 3"&gt;&lt;/img&gt;



 &lt;hr&gt;&lt;/hr&gt;



 &lt;h2&gt;WSL Comfort：顺手配置 WSL&lt;/h2&gt;



 &lt;p&gt;Windows Developer Config 还提供了一套   &lt;strong&gt;WSL Comfort&lt;/strong&gt; 工具，专门用来把刚装好的 WSL 命令行环境配置得更顺手。&lt;/p&gt;



 &lt;p&gt;只需要运行：&lt;/p&gt;


 &lt;div&gt;  &lt;pre&gt;
.\wsl-comfort\install.ps1
&lt;/pre&gt;&lt;/div&gt;


 &lt;p&gt;就会进入交互式设置，让你：&lt;/p&gt;



 &lt;ul&gt;
  &lt;li&gt;选择    &lt;strong&gt;zsh 或 bash&lt;/strong&gt;，&lt;/li&gt;



  &lt;li&gt;自动安装常用工具：   &lt;strong&gt;Starship、fzf、ripgrep、fd、bat、eza、zoxide、jq、Homebrew&lt;/strong&gt; 等&lt;/li&gt;



  &lt;li&gt;配置 Cascadia Code Nerd Font 和 Windows Terminal 的 WSL Profile&lt;/li&gt;
&lt;/ul&gt;



 &lt;p&gt;还支持   &lt;code&gt;-NonInteractive&lt;/code&gt; 无人值守安装。&lt;/p&gt;



 &lt;p&gt;另外，其中负责 Linux 环境配置的   &lt;code&gt;comfort-shell-bootstrap.sh&lt;/code&gt; 也可以脱离 WSL，单独用于普通的 Ubuntu 主机。&lt;/p&gt;



 &lt;p&gt;也就是说，你可以非常容易的将同一套终端环境，复用到多台设备上。&lt;/p&gt;



 &lt;hr&gt;&lt;/hr&gt;



 &lt;h2&gt;只装需要的开发环境&lt;/h2&gt;



 &lt;p&gt;对于不需要一键预装一大堆开发软件的用户来说，还有一个   &lt;strong&gt;Single-language workloads&lt;/strong&gt; 功能，可以按需配置某一种语言或开发环境。&lt;/p&gt;



 &lt;p&gt;目前包括   &lt;strong&gt;TypeScript、Python、Go、Rust、Java、PHP、.NET、SQL、PowerShell、WinForms、WinUI 3、Windows App CLI&lt;/strong&gt; 等。&lt;/p&gt;



 &lt;p&gt;例如，只需要 Python：&lt;/p&gt;


 &lt;div&gt;  &lt;pre&gt;
.\Workloads\python\install.ps1
&lt;/pre&gt;&lt;/div&gt;


 &lt;p&gt;就会自动安装   &lt;strong&gt;Python 3.14 + uv&lt;/strong&gt;。&lt;/p&gt;



 &lt;p&gt;只需要 Java：&lt;/p&gt;


 &lt;div&gt;  &lt;pre&gt;
.\Workloads\java\install.ps1
&lt;/pre&gt;&lt;/div&gt;


 &lt;p&gt;则会安装   &lt;strong&gt;Microsoft Build of OpenJDK 25 LTS&lt;/strong&gt;。&lt;/p&gt;



 &lt;p&gt;这样无论是准备一台完整的 Windows 开发机，还是只想快速搭一套 Python、Java、Rust 开发环境，都可以直接使用现成的配置。&lt;/p&gt;



 &lt;hr&gt;&lt;/hr&gt;



 &lt;h2&gt;获取&lt;/h2&gt;



 &lt;ul&gt;
  &lt;li&gt;   &lt;a href="https://github.com/microsoft/WindowsDeveloperConfig" rel="noopener" target="_blank"&gt;GitHub&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;



 &lt;p&gt;怎么样，是不是有一种微软秘密武器的感觉？&lt;/p&gt;
 &lt;hr&gt;&lt;/hr&gt; &lt;h2&gt;相关阅读&lt;/h2&gt; &lt;ul&gt;  &lt;li&gt;   &lt;a href="https://www.appinn.com/ultimate-list-of-free-windows-software-from-microsoft/" rel="bookmark" title="Permanent Link: &amp;#24494;&amp;#36719;&amp;#30340; 150 &amp;#27454;&amp;#20813;&amp;#36153;&amp;#36719;&amp;#20214;[&amp;#37096;&amp;#20998;&amp;#65292;&amp;#24453;&amp;#26356;&amp;#26032;]"&gt;微软的 150 款免费软件[部分，待更新]&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://www.appinn.com/ms-ai-for-beginners/" rel="bookmark" title="Permanent Link: &amp;#24494;&amp;#36719;&amp;#65306;&amp;#20154;&amp;#24037;&amp;#26234;&amp;#33021;&amp;#21021;&amp;#23398;&amp;#32773;&amp;#35838;&amp;#31243;&amp;#65292;&amp;#19968;&amp;#20849;12&amp;#21608;&amp;#12289;24&amp;#33410;&amp;#35838;"&gt;微软：人工智能初学者课程，一共12周、24节课&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://www.appinn.com/powertoys-v0-62-0/" rel="bookmark" title="Permanent Link: &amp;#24494;&amp;#36719;&amp;#23448;&amp;#26041;&amp;#36229;&amp;#23454;&amp;#29992; 15+ &amp;#23567;&amp;#24037;&amp;#20855;&amp;#38598; PowerToys v0.62.0 &amp;#21457;&amp;#24067;&amp;#65292;&amp;#26032;&amp;#22686;&amp;#25991;&amp;#26412;&amp;#25552;&amp;#21462;&amp;#22120; OCR &amp;#21151;&amp;#33021;"&gt;微软官方超实用 15+ 小工具集 PowerToys v0.62.0 发布，新增文本提取器 OCR 功能&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://www.appinn.com/win11debloat/" rel="bookmark" title="Permanent Link: Win11Debloat &amp;#20013;&amp;#25991;&amp;#29256; &amp;#8211; &amp;#24494;&amp;#36719;&amp;#27424;&amp;#25105;&amp;#30340;&amp;#24615;&amp;#33021;&amp;#35813;&amp;#36824;&amp;#20102;&amp;#65306;&amp;#19968;&amp;#38190;&amp;#21368;&amp;#36733; 90+ &amp;#27454; Windows 11 &amp;#39044;&amp;#35013;&amp;#36719;&amp;#20214;[2026.6.24&amp;#26356;&amp;#26032;]"&gt;Win11Debloat 中文版 – 微软欠我的性能该还了：一键卸载 90+ 款 Windows 11 预装软件[2026.6.24更新]&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://www.appinn.com/windows-defender-enable-microsoft-maps/" rel="bookmark" title="Permanent Link: &amp;#21152;&amp;#22266; Windows Defender &amp;#65292;&amp;#24320;&amp;#21551;&amp;#24494;&amp;#36719;&amp;#20113;&amp;#20445;&amp;#25252;&amp;#65292;&amp;#21033;&amp;#29992;&amp;#12300;&amp;#24494;&amp;#36719;&amp;#39640;&amp;#32423;&amp;#20445;&amp;#25252;&amp;#26381;&amp;#21153;&amp;#12301;&amp;#65288;MAPS&amp;#65289;&amp;#26469;&amp;#23454;&amp;#26102;&amp;#39044;&amp;#38450;&amp;#26410;&amp;#30693;&amp;#30149;&amp;#27602;"&gt;加固 Windows Defender ，开启微软云保护，利用「微软高级保护服务」（MAPS）来实时预防未知病毒&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt; &lt;hr&gt;&lt;/hr&gt;
 &lt;a href="http://www.appinn.com/copyright/?utm_source=feeds&amp;utm_medium=copyright&amp;utm_campaign=feeds" title="&amp;#29256;&amp;#26435;&amp;#22768;&amp;#26126;"&gt;©&lt;/a&gt;2021 青小蛙 for  &lt;a href="http://www.appinn.com/?utm_source=feeds&amp;utm_medium=appinn&amp;utm_campaign=feeds" title="&amp;#26412;&amp;#25991;&amp;#26469;&amp;#33258;&amp;#23567;&amp;#20247;&amp;#36719;&amp;#20214;"&gt;小众软件&lt;/a&gt; |  &lt;a href="http://www.appinn.com/join-us/?utm_source=feeds&amp;utm_medium=joinus&amp;utm_campaign=feeds" title="&amp;#21152;&amp;#20837;&amp;#23567;&amp;#20247;&amp;#36719;&amp;#20214;"&gt;加入我们&lt;/a&gt; |  &lt;a href="https://meta.appinn.net/c/faxian/?utm_source=feeds&amp;utm_medium=contribute&amp;utm_campaign=feeds" rel="noopener" target="_blank" title="&amp;#32473;&amp;#23567;&amp;#20247;&amp;#36719;&amp;#20214;&amp;#25237;&amp;#31295;"&gt;投稿&lt;/a&gt; |  &lt;a href="http://www.appinn.com/feeds-subscribe/?utm_source=feeds&amp;utm_medium=feedsubscribe&amp;utm_campaign=feeds" target="_blank" title="&amp;#21487;&amp;#20197;&amp;#20998;&amp;#31867;&amp;#35746;&amp;#38405;&amp;#23567;&amp;#20247;&amp;#65292;Windows/MAC/&amp;#28216;&amp;#25103;"&gt;订阅指南&lt;/a&gt; &lt;br /&gt; 3659b075e72a5b7b1b87ea74aa7932ff  &lt;br /&gt;
 &lt;a href="https://www.appinn.com/microsoft-windows-developer-config/#comments" title="to the comments"&gt;点击这里留言、和原作者一起评论&lt;/a&gt; &lt;p&gt;  &lt;a href="https://www.appinn.com/microsoft-windows-developer-config/"&gt;[ 点击前往获取链接 ]&lt;/a&gt;&lt;/p&gt; &lt;hr&gt;&lt;/hr&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>Windows 开发环境 开发者 开发者工具</category>
      <guid isPermaLink="true">https://itindex.net/detail/63288-%E5%BE%AE%E8%BD%AF-%E5%AE%98%E6%96%B9-%E5%BC%80%E6%BA%90</guid>
      <pubDate>Sun, 27 Sep 2026 20:13:10 CST</pubDate>
    </item>
    <item>
      <title>开放权重模型逐渐形成了生态-软硬件、量化、Agent、推理引擎、后训练</title>
      <link>https://itindex.net/detail/63287-%E5%BC%80%E6%94%BE-%E6%9D%83%E9%87%8D-%E6%A8%A1%E5%9E%8B</link>
      <description>&lt;div&gt;    &lt;div&gt;      &lt;div&gt;        &lt;div&gt;          &lt;h2&gt;开放权重模型的竞争，已经不只是比谁 Benchmark 更高了。

Nathan Lambert 最新这篇文章，把能力、价格、下载量、真实调用和学术采用放到了一起看。

一个很明显的变化是，中国开放权重模型已经形成了自己的生态。

按他的统计，中国模型在 Hugging Face 的累计下载量约 32 亿，是美国的两倍；OpenRouter 上开放模型每周使用量已经从一年前约 1T Token 涨到 80T，其中中国模型占比超过 80%。

学术圈也很明显。Qwen 现在出现在约 30% 的 AI 论文里，中国开放权重模型整体已经超过 40%。

更关键的是，Lambert 估计最强的中国开放权重模型距离美国闭源前沿只差大约 2–5 个月。这当然是他的估计，但已经足够说明变化有多快。

所以我现在更在意的不是「开放模型能不能追上闭源」。

而是谁的模型会成为别人做研究、搭 Agent、做产品时默认踩着的那一层。

一旦形成这种习惯，影响的就是整个开发生态。

文章：            &lt;a href="https://www.interconnects.ai/p/the-current-balance-of-power-in-open" rel="noopener noreferrer" target="_blank"&gt;interconnects.ai/p/the-current-…&lt;/a&gt;&lt;/h2&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;  &lt;div&gt;        &lt;div&gt;          &lt;div&gt;            &lt;div&gt;              &lt;div&gt;                &lt;a href="https://x.com/Xudong07452910/status/2102677612841587124/photo/1"&gt;                  &lt;img alt="" height="906" src="https://pbs.twimg.com/media/HS42AGQW8AAjova?format=webp&amp;name=large" width="1116"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/div&gt;              &lt;a href="https://x.com/Xudong07452910/status/2102677612841587124/photo/1"&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;        &lt;div&gt;    &lt;br /&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63287-%E5%BC%80%E6%94%BE-%E6%9D%83%E9%87%8D-%E6%A8%A1%E5%9E%8B</guid>
      <pubDate>Thu, 24 Sep 2026 11:43:22 CST</pubDate>
    </item>
    <item>
      <title>不要被 AI 炒作愚弄</title>
      <link>https://itindex.net/detail/63286-ai-%E7%82%92%E4%BD%9C</link>
      <description>Anthropic 声称其模型 Claude Mythos 在发现软件漏洞上胜过大多数安全专家。随后发生了 OpenAI–Hugging Face 安全事件，此后 Anthropic（自豪）和 Meta（不情愿）也披露了各自模型的类似事件。紧接着 Anthropic 宣称其模型取得了数学领域的突破；OpenAI 也声称自己取得了数学突破。Anthropic 工程师 Jacob Coxon 在宣布离职时引发了广泛关注，他声称该公司与 OpenAI 正“冲向自我进化的超级智能，并拿我们的生命在赌博”。媒体大肆报道了这些事件，且沿用了相关公司赋予其软件的拟人化叙事——即把软件描绘成不仅功能强大，而且已初具通用人工智能（AGI）雏形的产物。但深入研究的专家则给出了不同的答案，虽然这些发现并不能吸引眼球。网络安全专家指出，涉及模型的安全事件更多是 OpenAI 的疏忽大意，未能采取基本的安全措施，而不是“模型失控”或“AI 智能体创造文明”。OpenAI 模型在解决数学难题上的突破其原创性也相当可疑。数学家公开对 AI 企业利用其专业领域进行炒作提出了警告。AI 公司通过炒作模型失控也将自己置身事外，将责任归咎于大模型而不是公司本身，逃避应承担的责任。以 OpenAI 为例，当该公司开发的恶意软件被用于入侵另一家公司时，媒体、名人和议员谈论是“失控模型”而不是 OpenAI 的责任，仿佛大模型真的会自动发动攻击，公众的注意力被转移到虚构的“超级智能”的恐惧之上。我们不要被 AI 公司的炒作所愚弄。
 &lt;p&gt;&lt;/p&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63286-ai-%E7%82%92%E4%BD%9C</guid>
      <pubDate>Wed, 23 Sep 2026 23:29:02 CST</pubDate>
    </item>
    <item>
      <title>Jev 工程学：为 coding agent 而作</title>
      <link>https://itindex.net/detail/63285-jev-%E5%B7%A5%E7%A8%8B%E5%AD%A6-coding</link>
      <description>&lt;div&gt;    &lt;h1&gt;Jev 工程学：为 coding agent 而作&lt;/h1&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#jev-&amp;#24037;&amp;#31243;&amp;#23398;&amp;#20026;-coding-agent-&amp;#32780;&amp;#20316;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;blockquote&gt;    &lt;p&gt;      &lt;strong&gt;2026 年 Jev 工程实践工作笔记&lt;/strong&gt;副标题：TypeSafe 创始人的 Jev 构建蓝图
A Synthesis for Study · Based on design notes by Diogo Almeida (TypeSafe)
独立汇编，2026 年 9 月。与 TypeSafe 无隶属关系，亦未获其背书。&lt;/p&gt;    &lt;p&gt;本文是      &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/assets/original.pdf"&gt;原文 PDF&lt;/a&gt;的完整中文翻译，保留原有结构与全部七张插图。&lt;/p&gt;&lt;/blockquote&gt;  &lt;blockquote&gt;    &lt;p&gt;      &lt;strong&gt;本仓库现在有两份文档&lt;/strong&gt;&lt;/p&gt;    &lt;ul&gt;      &lt;li&gt;本页（根目录）：第三方根据笔记        &lt;strong&gt;汇编整理&lt;/strong&gt;的工作笔记——结构正式、有图表、「六个症状」已表格化&lt;/li&gt;      &lt;li&gt;        &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/source-notes"&gt;          &lt;code&gt;source-notes/&lt;/code&gt;&lt;/a&gt;：作者的        &lt;strong&gt;原始设计笔记&lt;/strong&gt;——保留了原始项目符号结构、随手记的链接，以及尚未定型的口语化表达&lt;/li&gt;&lt;/ul&gt;    &lt;p&gt;两份对照读，能看到从笔记到成文的加工过程。原始笔记里有汇编版没有的内容（如 SLOP 子任务拆解、附录 2 引用的原始推文、以及作者的推理过程）。&lt;/p&gt;&lt;/blockquote&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig1.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 1&amp;#65306;Jev harness" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig1.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 1. Jev harness。显式状态（左）以可寻址、带类型的块存储。Jev（右）回答每一轮的问题：显示哪些块、是否复用缓存、用哪个模型和工具、以及一条命令能否执行。上下文组装层与路由器驱动运行时，运行时以「片段优先」的方式披露数百个工具，并在安全策略下把工作路由到前沿模型、子 agent、廉价模型或后台模型。&lt;/em&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;摘要&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#&amp;#25688;&amp;#35201;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;Coding agent 简单得出人意料。大多数就是一个模型加少量工具的 while 循环，而模型本身之外几乎没有什么有意义的创新。&lt;/p&gt;  &lt;p&gt;这份笔记综合了 TypeSafe 创始人 Diogo Almeida 的设计笔记，讲的是围绕 Jev 构建 coding agent 的思路——Jev 是 TypeSafe 的决策模型，把结构化状态转换成类型化输出，例如 choice、score 和 noul 决策。&lt;/p&gt;  &lt;p&gt;笔记从一个挑衅性的问题开始：    &lt;strong&gt;如果语言模型没有 KV cache，你会怎么设计一个 coding agent？&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;这个问题暴露了当前 agent 不加审视就继承的六个设计选择：    &lt;strong&gt;亏损的路由、挤占上下文的工具调用、盲目压缩的 compaction、几乎不触发的子 agent、丢弃状态的重启、以及电池之争。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;我们呈现笔记提出的替代架构：一个围绕显式、带类型状态构建的 harness，由 Jev 做每一轮的决策。其中上下文块按查询逐个打分，路由定价时计入重新处理的成本，工具分级披露，指令按条件加载，只读的后台任务共享一次检索。&lt;/p&gt;  &lt;p&gt;我们还演算一遍路由算术、真实 agent 会话的 token 经济学，以及按文件敏感度（而不仅是难度）路由的安全论证。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;索引词&lt;/strong&gt;：Jev、TypeSafe、coding agent、harness 工程、KV cache、上下文工程、模型路由、子 agent、工具调用、compaction、AGENTS.md、后台 agent。&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;I. 为什么还要再做一个 agent&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#i-&amp;#20026;&amp;#20160;&amp;#20040;&amp;#36824;&amp;#35201;&amp;#20877;&amp;#20570;&amp;#19968;&amp;#20010;-agent"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;笔记以四个工作假设开场。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第一，coding agent 很简单&lt;/strong&gt;，尤其是 agent 的部分：一个循环、一个模型、少量工具。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第二，现有 agent 中最好的部分可以复用。&lt;/strong&gt;前沿模型都能通过 API 拿到，开源项目提供了灵感和甚至 UI 组件，而只有少数领域（比如 MCP 服务器的认证）才带有真正的复杂性。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第三，第一方 agent 的成本优势可能在缩小&lt;/strong&gt;，因为用量正从捆绑订阅转向按量付费的 API 定价。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第四，也是最重要的，有些能力只能原生构建。&lt;/strong&gt;它们无法作为插件交付给别人的 agent，因为它们要求控制    &lt;strong&gt;每一轮上下文是如何被组装的&lt;/strong&gt;。&lt;/p&gt;  &lt;p&gt;第四个假设就是论点。    &lt;strong&gt;如果 agent 只是一个循环，那么杠杆不在循环里。杠杆在于循环每一轮喂给模型的东西。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;A. Jev 站在哪里&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-jev-&amp;#31449;&amp;#22312;&amp;#21738;&amp;#37324;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;    &lt;strong&gt;Jev 不是写代码的模型，它是旁边那层决策层。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;harness 把当前应用状态（目标、上下文、规则、可用动作、之前的动作）和一个预定义问题交给 Jev，Jev 返回一个类型化答案：choice、score 或 noul，每个都带概率。然后由前沿模型、子 agent、工具和确定性代码去做实际工作。&lt;/p&gt;  &lt;p&gt;因为输出是类型化的而不是自由文本，harness 可以验证它、施加阈值、基于它分支，而不必解析散文。&lt;/p&gt;  &lt;p&gt;这份笔记里每一个原生功能，底层都是在高频决策点上问 Jev 一个问题。    &lt;strong&gt;每个会话被问上千次，那些小决策才是杠杆所在。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;B. 组织起一切的那个问题&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#b-&amp;#32452;&amp;#32455;&amp;#36215;&amp;#19968;&amp;#20999;&amp;#30340;&amp;#37027;&amp;#20010;&amp;#38382;&amp;#39064;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;笔记描述了一个他喜欢问其他工程师的问题：    &lt;strong&gt;如果语言模型没有 KV cache，你会怎么设计 coding agent？&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;KV cache 是 agent 被建成「只追加对话记录」的原因。复用缓存前缀很便宜；改动上下文早期的东西会让缓存失效，迫使模型重新处理改动之后的一切。    &lt;strong&gt;这个单一的经济事实塑造了当前 agent 的几乎每一个设计决策——而且通常没人把这件事说出来。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;想象把缓存拿掉，会做两件事。它允许一种为 Jev 设计的、状态显式的架构——    &lt;strong&gt;上下文是被组装的，而不是被累积的&lt;/strong&gt;。同时它揭示为什么「把简单活路由给便宜模型」这种直觉上正确的想法在实践中会失败。&lt;/p&gt;  &lt;p&gt;笔记管这叫    &lt;strong&gt;「KV cache 的暴政」&lt;/strong&gt;。&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;II. KV cache 的六个症状&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#ii-kv-cache-&amp;#30340;&amp;#20845;&amp;#20010;&amp;#30151;&amp;#29366;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;table&gt;    &lt;tr&gt;      &lt;th&gt;症状&lt;/th&gt;      &lt;th&gt;为什么存在&lt;/th&gt;      &lt;th&gt;代价&lt;/th&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;1. 路由失效&lt;/td&gt;      &lt;td&gt;交还给大模型时需要重新处理上下文&lt;/td&gt;      &lt;td&gt;混合路由比纯前沿更贵&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;2. 工具挤占上下文&lt;/td&gt;      &lt;td&gt;schema 必须待在系统消息里&lt;/td&gt;      &lt;td&gt;token 花在无关工具上；选择质量下降&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;3. Compaction&lt;/td&gt;      &lt;td&gt;假设所有未来轮次共享同一份状态&lt;/td&gt;      &lt;td&gt;查询无关的压缩丢掉了后面需要的东西&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;4. 子 agent 稀少&lt;/td&gt;      &lt;td&gt;决定传什么上下文进去、什么合并回来很难&lt;/td&gt;      &lt;td&gt;几乎没有自动并行&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;5. 重启&lt;/td&gt;      &lt;td&gt;有状态的对话记录会随时间腐坏&lt;/td&gt;      &lt;td&gt;相关的旧状态跟着坏状态一起被丢弃&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;6. 电池之争&lt;/td&gt;      &lt;td&gt;每个内置能力都永久消耗上下文&lt;/td&gt;      &lt;td&gt;在易用与强大之间被迫二选一&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;  &lt;div&gt;    &lt;h3&gt;A. 路由不工作&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-&amp;#36335;&amp;#30001;&amp;#19981;&amp;#24037;&amp;#20316;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;直觉方案是让前沿模型规划、把执行交给便宜模型、再让前沿模型回来审查。笔记用 Opus（$5 输入 / $25 输出每百万 token）和 Sonnet（$3 / $15）的标价算了这笔账。设 X 是上下文 token，Y 是生成的输出 token，Z 是工作中额外读入的 token（比如命令输出和文件读取）。&lt;/p&gt;  &lt;div&gt;    &lt;pre&gt;      &lt;code&gt;路径 1  纯 Opus
  生成        25·Y
  读取        5·Z
  合计        25Y + 5Z

路径 2  Opus → Sonnet → Opus
  Sonnet 载入上下文         3·X
  Sonnet 生成               15·Y
  Sonnet 读取               3·Z
  Opus 重新加载变化的部分    5·(Y+Z)
  合计                      3X + 20Y + 8Z&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;  &lt;p&gt;当会话很长（X 大）、或工作读取远多于写入、或两者兼有时，路径 2 更贵。    &lt;strong&gt;代入一个合理的会话形态 X = 0.65、Y = 0.12、Z = 0.23，纯 Opus 是 4.15，路由路径是 6.19。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;留在前沿模型上的成本，大约是那条本该省钱的路径的三分之二。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig2.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 2&amp;#65306;&amp;#36335;&amp;#30001;&amp;#38519;&amp;#38449;" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig2.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 2. 路由陷阱。把执行委派给便宜模型，会在下行时加一次上下文载入、在回程时加一次重新处理，两者加起来超过了每 token 的折扣。&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;教训不是「路由是错的」，而是    &lt;strong&gt;按 token 定价、而不是按上下文重建定价的路由是错的&lt;/strong&gt;。路由只有在两个条件下才可行：    &lt;strong&gt;harness 能给便宜模型一个小而专用的上下文（而不是完整记录），且回程不强迫前沿模型重读助手产出的一切。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;B. 工具调用是个奇怪的取舍&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#b-&amp;#24037;&amp;#20855;&amp;#35843;&amp;#29992;&amp;#26159;&amp;#20010;&amp;#22855;&amp;#24618;&amp;#30340;&amp;#21462;&amp;#33293;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;工具必须提前在系统消息里声明，带着完整的参数 schema，不管这轮用不用得上。这消耗大量上下文，而且笔记的判断是：    &lt;strong&gt;它并不产生特别聪明的工具选择。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;工作假设是：模型在「高基数」（一次给太多工具）和「离策略工具调用」（工具的使用模式与训练时见过的不符）的某种组合上挣扎。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;这可能就是为什么 skills——加载一句短描述、细节延后——常常胜过裸工具列表和 MCP 服务器。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;C. Compaction 存在&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#c-compaction-&amp;#23384;&amp;#22312;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;如果每个未来轮次都想要同一份共享状态，compaction 就完全合理。笔记质疑的正是这个假设。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;Compaction 试图做通用压缩，这既难又有损。而查询感知的压缩容易得多：如果你知道下一个问题是什么，你就知道该留什么。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;在问题已知之前写下的摘要，一定会丢掉问题需要的东西。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;D. 子 agent 很平庸&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#d-&amp;#23376;-agent-&amp;#24456;&amp;#24179;&amp;#24248;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;模型并行的程度低于预期。怀疑的原因在状态管理：    &lt;strong&gt;决定父上下文的哪些部分要传进去、每个子 agent 的发现中哪些要合并回来。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;当这个决策既昂贵又容易出错时，模型就会回避它。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;E. 重启存在&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#e-&amp;#37325;&amp;#21551;&amp;#23384;&amp;#22312;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;重启会话是对话记录漂移或损坏时的标准补救。    &lt;strong&gt;但它把坏状态和好状态一起丢掉了。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;有了可寻址状态，替代方案是干净启动、按需只重新加载仍然相关的旧块。&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;F. 电池并非内置&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#f-&amp;#30005;&amp;#27744;&amp;#24182;&amp;#38750;&amp;#20869;&amp;#32622;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;关于 agent 是否该内置能力，有一场持续的争论。今天它是在「易用」（重度内置的 agent 在一端）和「强大」（Claude Code、Codex 这类高级用户工具在另一端）之间的取舍。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;这个取舍存在，仅仅因为每块电池都永久消耗上下文。把那个成本拿掉，争论就消解了。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;III. token 到底花在哪&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#iii-token-&amp;#21040;&amp;#24213;&amp;#33457;&amp;#22312;&amp;#21738;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;在重新设计 harness 之前，先搞清楚一次会话里哪个部分吃掉了预算。下表估算典型 CLI coding agent 会话中，各子任务占总处理 token 的份额。    &lt;strong&gt;这是输入密集视角，所以重复读取每次计数。&lt;/strong&gt;&lt;/p&gt;  &lt;table&gt;    &lt;tr&gt;      &lt;th&gt;子任务&lt;/th&gt;      &lt;th&gt;约占&lt;/th&gt;      &lt;th&gt;说明&lt;/th&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;读文件内容&lt;/td&gt;      &lt;td&gt;30–40%&lt;/td&gt;      &lt;td&gt;最大项；文件作为上下文被反复读取&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;搜索代码库&lt;/td&gt;      &lt;td&gt;10–18%&lt;/td&gt;      &lt;td&gt;grep、glob、列表；输出噪音大&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;命令输出&lt;/td&gt;      &lt;td&gt;10–20%&lt;/td&gt;      &lt;td&gt;堆栈和日志在失败时爆开&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;系统提示、工具 schema、AGENTS.md&lt;/td&gt;      &lt;td&gt;5–12%&lt;/td&gt;      &lt;td&gt;每一轮都付的固定开销&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;对话重放放大器&lt;/td&gt;      &lt;td&gt;—&lt;/td&gt;      &lt;td&gt;为什么上面每一项都被重复计数&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;推理与规划&lt;/td&gt;      &lt;td&gt;5–15%&lt;/td&gt;      &lt;td&gt;困难调试时更高&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;写和编辑代码&lt;/td&gt;      &lt;td&gt;4–10%&lt;/td&gt;      &lt;td&gt;diff 和 str_replace 编辑很紧凑&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;向用户解释&lt;/td&gt;      &lt;td&gt;2–5%&lt;/td&gt;      &lt;td&gt;CLI agent 默认简洁&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig3.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 3&amp;#65306;&amp;#26816;&amp;#32034;&amp;#21344;&amp;#20027;&amp;#23548;" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig3.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 3. 检索占主导。读取、搜索和命令输出合计约占处理 token 的三分之二；写代码不到十分之一。&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;最触目惊心的是倒数第二行：写代码，也就是 coding agent 存在的理由，是其中最小的开支项之一。读取和搜索占主导。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;独立的分析指向同一方向：    &lt;strong&gt;微软的 fastcontext 项目报告，在 GPT-5.4 的轨迹里，读取和搜索占所有工具调用轮次的 56.2%、占主 agent 总 token 的 46.5%。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;如果这能推广，coding agent 里最大的效率收益不是更好的模型或更好的 diff 格式，而是更聪明的检索。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;IV. 基础层：权限与工具路由&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#iv-&amp;#22522;&amp;#30784;&amp;#23618;&amp;#26435;&amp;#38480;&amp;#19982;&amp;#24037;&amp;#20855;&amp;#36335;&amp;#30001;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;有两项改进可以集成到任何现有 agent 里，无论原生与否。&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;A. 可编程权限&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-&amp;#21487;&amp;#32534;&amp;#31243;&amp;#26435;&amp;#38480;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;每一条 agent 运行的命令都提出「它到底该不该跑」的问题。Claude 的 auto 模式用分类器回答这个。    &lt;strong&gt;笔记提议走得更远：权限表达为对「什么允许、什么不允许」的可编程查询，&lt;/strong&gt;而且在风险足够高时做更深的检查——比如在执行前读取 Python 或 shell 文件的内容，而不是只批准命令名。&lt;/p&gt;  &lt;div&gt;    &lt;pre&gt;policy &amp;quot;exec&amp;quot;:deny   if command touches ~/.ssh or .env*deny   if script contents contain network egressand task.scope != &amp;quot;deploy&amp;quot;ask    if command writes outside repo rootallow  if command in read_only_setallow  if tests/ and exit code is expected&lt;/pre&gt;&lt;/div&gt;  &lt;div&gt;    &lt;h3&gt;B. harness 作为工具路由器&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#b-harness-&amp;#20316;&amp;#20026;&amp;#24037;&amp;#20855;&amp;#36335;&amp;#30001;&amp;#22120;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;不是把每个工具 schema 都暴露给模型，而是让 harness 坐在意图和调用之间。    &lt;strong&gt;模型用纯文本描述它想干什么。&lt;/strong&gt;harness 然后用一串类型化的 Jev 调用来选出最合适的那个工具（或前几名候选），并构造参数。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;模型永远不必在上下文里持有几百个 schema，而错误的参数类型会变成验证错误，而不是静默失败。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;V. 元注意力：把上下文本身当成决策&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#v-&amp;#20803;&amp;#27880;&amp;#24847;&amp;#21147;&amp;#25226;&amp;#19978;&amp;#19979;&amp;#25991;&amp;#26412;&amp;#36523;&amp;#24403;&amp;#25104;&amp;#20915;&amp;#31574;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;核心提议去掉了「上下文是静态的」这个观念。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;对每一个用户查询，harness 问 Jev 两件事：&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第一，之前的上下文有多好&lt;/strong&gt;——复用现有 KV cache 是对的吗，还是从头重建更便宜也更好？    &lt;strong&gt;这被框定为一个显式的、成本感知的决策，而不是默认行为。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第二，如何构造一个包含所有相关的、且不含任何多余的新上下文。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;最简单的形态是：对每一个上下文块做一次 noul——每次工具调用的输入、每次工具调用的输出、每一段内部推理、以及可能每一次与用户的交流。    &lt;strong&gt;后续版本把分数换成可见性级别。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig4.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 4&amp;#65306;&amp;#21487;&amp;#35265;&amp;#24615;&amp;#38454;&amp;#26799;" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig4.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 4. 可见性阶梯。同一个块可以被隐藏、简短总结、详细总结或完整展示，取决于当前查询。      &lt;strong&gt;这就是查询感知的压缩——compaction 所缺的那个属性。&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;收益是：    &lt;strong&gt;compaction 背后的想法活下来了，但它的主要缺陷消失了。Compaction 在知道问题之前压缩一次；阶梯在知道问题之后按查询压缩。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;一个 2400 行的 grep 结果，对某个问题可以是十二条相关命中，对下一个问题可以是不可见的——而它从未从状态里被删除。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;笔记还加了一个值得注意的视觉想法：    &lt;strong&gt;如果 harness 能热力图显示 grep 输出里哪部分相关，它就能按预算允许的程度随意过滤那份输出。&lt;/strong&gt;而且，笔记观察到，    &lt;strong&gt;它看起来也会很酷。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;VI. 路由与子 agent 重访&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#vi-&amp;#36335;&amp;#30001;&amp;#19982;&amp;#23376;-agent-&amp;#37325;&amp;#35775;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;    &lt;strong&gt;一等公民的动态上下文，是让路由重新可行的原因。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;一旦 harness 能为子任务构造一个小而相关的上下文，把这个子任务交给更便宜或更快的模型就不需要便宜模型载入整个会话，    &lt;strong&gt;而结果可以作为被评分的块合并回来，而不是一份前沿模型必须重读的记录。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;同一机制解锁了子 agent。    &lt;strong&gt;笔记的猜测是：今天子 agent 成本的大部分，是决定传什么上下文的工作&lt;/strong&gt;——相比之下，用户敲一条简单指令反而容易。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;如果构造那个上下文变得便宜且自动，子 agent 就能用得频繁得多。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;它还打开了一个面向用户的控制：    &lt;strong&gt;多花钱换更快或更好的结果，还是保守运行、尽量少花。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;A. 极端并行&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-&amp;#26497;&amp;#31471;&amp;#24182;&amp;#34892;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;如果派发任务变得便宜，很多任务会同时跑，harness 就继承了并发系统的全部问题：    &lt;strong&gt;同步原语、agent 间通信、多个 agent 共享状态时的写冲突。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;笔记建议用带锁的共享状态。    &lt;strong&gt;而显式区分读与写让它变得可处理，因为只读任务永远不会争锁。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;B. 目标去重&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#b-&amp;#30446;&amp;#26631;&amp;#21435;&amp;#37325;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;对 /goal 这类目标驱动的循环，一个悬而未决的问题是：agent 会不会重复劳动。&lt;/p&gt;  &lt;p&gt;一个缓解措施：    &lt;strong&gt;在派发任何子任务之前，把它注册为子目标，并与所有历史子目标去重。已经做过的、或正在飞行中的工作，永远不会被启动两次。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;VII. 工具与技能，从第一性原理出发&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#vii-&amp;#24037;&amp;#20855;&amp;#19982;&amp;#25216;&amp;#33021;&amp;#20174;&amp;#31532;&amp;#19968;&amp;#24615;&amp;#21407;&amp;#29702;&amp;#20986;&amp;#21457;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;笔记主张在「今天总是加载的工具」和「按需加载的 skills」之间加一层。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;模型无法提出一个它不知道存在的动作&lt;/strong&gt;，所以它需要短片段来描述有什么可用，类似于 skill 描述。    &lt;strong&gt;但这些片段不必住在系统消息里——它们可以在相关时动态加载。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;它们背后是「在需要时倾倒出可用动作的完整 schema」的能力，类似一个工具搜索工具。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;把这一切绑在一起的要求是：这些东西在不再需要之后，都不能污染上下文。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig5.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 5&amp;#65306;&amp;#19977;&amp;#32423;&amp;#25259;&amp;#38706;" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig5.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 5. 三级披露。模型看到一张廉价的全局地图，只为它选中的东西付细节的钱，用完之后把细节从上下文中丢掉。&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;如果这能成，第二节里的电池之争就消失了：当一个内置能力在被使用之前几乎不花成本，agent 就能带着几百个工具和几千页文档发布。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;笔记还点出第二个好处：    &lt;strong&gt;近乎零成本的集成是强有力的联合营销渠道，而且它们让事情「就是能用」。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;A. 更好的电池&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-&amp;#26356;&amp;#22909;&amp;#30340;&amp;#30005;&amp;#27744;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;很多社区工具承诺帮助 agent，实际上并没有。笔记举了一个工具输出压缩器作为例子，并猜测原因：    &lt;strong&gt;模型并不原生理解这些工具。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;一个原生 harness 可以发布第一方的提示词，教模型如何使用每一个工具&lt;/strong&gt;——实际上相当于每个工具配一个内置的子 agent 或 skill。而它干净的上下文意味着工具的自定义逻辑不会污染会话的其余部分。&lt;/p&gt;  &lt;p&gt;这里也有营销角度：    &lt;strong&gt;为你当前流行的工具发布集成，能让 agent 保持在对话之中。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;VIII. 条件化指令&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#viii-&amp;#26465;&amp;#20214;&amp;#21270;&amp;#25351;&amp;#20196;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;    &lt;strong&gt;今天的 AGENTS.md 是全量加载的。笔记提议按条件加载。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;改前端代码就加载风格指南；在某个子目录里工作就加载那个目录的坑文件（    &lt;strong&gt;而且笔记建议每个子目录都该有一个&lt;/strong&gt;）。&lt;/p&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig6.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 6&amp;#65306;&amp;#26465;&amp;#20214;&amp;#21270; AGENTS.md" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig6.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 6. 条件化 AGENTS.md。指令附着到条件上而不是会话上，而且条件片段会被钉住，使 compaction 无法把它摘要掉。&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;这像 skills，但笔记画出了一个区别：    &lt;strong&gt;skills 通常意味着「现在做这个」。条件化指令意味着「把这段记在某处」。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;第二种还需要一个 skills 没有的属性：免疫于压缩。&lt;/strong&gt;一个在长会话早期加载的 skill 最终会被压缩或摘要掉；    &lt;strong&gt;而一个绑定到条件的指令会在条件成立时被重新加载。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;笔记在普通聊天里观察到同样的需求：请求摘要应该拉入一个偏好的格式；请求用自己风格写作应该拉入样本和一份「不要什么」的清单；请求代码应该拉入风格指南和一条「不要加一百万条断言」的提示。&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;A. 结构化 skills&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#a-&amp;#32467;&amp;#26500;&amp;#21270;-skills"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;取决于 harness 变得多可编程，skills 可以携带行为变更而不只是指令，类似 Claude Code 的 skill hooks 但更强大。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;笔记指出的当前 hook 系统的一个限制是：一旦加上，hook 就永久留在会话里。结构化 skills 会随触发它们的条件附着和脱离行为。&lt;/strong&gt;&lt;/p&gt;  &lt;div&gt;    &lt;h3&gt;B. 递归语言模型&lt;/h3&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#b-&amp;#36882;&amp;#24402;&amp;#35821;&amp;#35328;&amp;#27169;&amp;#22411;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;一个相关方向是递归语言模型（recursive language models）的方法，它把 agent 更多的状态当作显式变量而不是对话记录文本。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;一个状态保存在具名变量里的世界，可能比「状态就是上下文窗口里恰好剩下的那些东西」的世界干净得多。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;IX. 安全感知的路由&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#ix-&amp;#23433;&amp;#20840;&amp;#24863;&amp;#30693;&amp;#30340;&amp;#36335;&amp;#30001;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;今天的路由围绕难度和成本。    &lt;strong&gt;笔记加了第三个轴：信任。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;一些通过低价供应商提供的开放权重模型比前沿 API 便宜得多，而笔记提出的担忧是：    &lt;strong&gt;经过某些端点的数据可能不再私有。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;提议是给每个子任务一个「可能触碰哪些文件」的估计，把策略附着到文件类型上，并据此路由。&lt;/p&gt;  &lt;table&gt;    &lt;tr&gt;      &lt;th&gt;可能触碰的文件&lt;/th&gt;      &lt;th&gt;策略&lt;/th&gt;      &lt;th&gt;可用模型&lt;/th&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;公开文档、开源依赖&lt;/td&gt;      &lt;td&gt;开放&lt;/td&gt;      &lt;td&gt;任意，最便宜的优先&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;应用代码&lt;/td&gt;      &lt;td&gt;标准&lt;/td&gt;      &lt;td&gt;经过审查的供应商&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;密钥、env、基础设施配置&lt;/td&gt;      &lt;td&gt;受限&lt;/td&gt;      &lt;td&gt;只用第一方前沿模型&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;专有研究代码&lt;/td&gt;      &lt;td&gt;自定义&lt;/td&gt;      &lt;td&gt;排除指定供应商&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;  &lt;p&gt;最后一行反映了一个更广的观点：    &lt;strong&gt;难度和成本不是路由的唯一理由。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;一个团队可能因为自己在做模型研究而避开某家供应商的模型，或因为安全工作而避开某些供应商。    &lt;strong&gt;一旦路由是策略驱动的，这些偏好就变成配置而不是纪律。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;X. 后台处理&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#x-&amp;#21518;&amp;#21488;&amp;#22788;&amp;#29702;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;笔记识别出几种流行 agent 工作流里的一个共同模式：&lt;/p&gt;  &lt;ul&gt;    &lt;li&gt;构建随工作进展而并行更新的 HTML 页面&lt;/li&gt;    &lt;li&gt;「理解而非生成才是新瓶颈」的论点&lt;/li&gt;    &lt;li&gt;在后台生成 eval&lt;/li&gt;    &lt;li&gt;用几张图加极少的字解释系统的 ELI5 技能&lt;/li&gt;    &lt;li&gt;让 agent 维护一个小型已部署的进度页面，带截图和笔记，      &lt;strong&gt;可以在长任务期间从手机查看&lt;/strong&gt;&lt;/li&gt;&lt;/ul&gt;  &lt;p&gt;还有一个相关的生产模式：    &lt;strong&gt;把线上流量镜像到候选模型，自动生成约一天的 eval，然后才决定是否切换。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;它们的共同点是：    &lt;strong&gt;在后台运行、作为正常工作流的扩展、而且都是当前代码库状态的只读函数。&lt;/strong&gt;跨模型审查（让一家供应商的 agent 审查另一家的产出）属于同一模式。&lt;/p&gt;  &lt;p&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/figures/fig7.png" rel="noopener noreferrer" target="_blank"&gt;      &lt;img alt="&amp;#22270; 7&amp;#65306;&amp;#26174;&amp;#24335;&amp;#29366;&amp;#24577;&amp;#19978;&amp;#30340;&amp;#21518;&amp;#21488;&amp;#22788;&amp;#29702;" src="https://github.com/yibie/jev-engineering-zh/raw/main/figures/fig7.png"&gt;&lt;/img&gt;&lt;/a&gt;&lt;/p&gt;  &lt;p&gt;    &lt;em&gt;图 7. 显式状态上的后台处理。找出与某个改动相关什么的这项昂贵工作，只做一次，由每个只读后台任务共享。&lt;/em&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;这正是 Jev 论点兑现的地方。&lt;/strong&gt;以 Jev 为中心的 harness 必须精确知道上下文里有什么、以及每个操作是读还是写。    &lt;strong&gt;找出与某个代码改动相关的信息，是件不平凡的工作。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;如果这份检索在所有后台任务间共享、而不是每个任务重复一遍，运行它们就会便宜得多，也就能经济地跑更多。&lt;/strong&gt;考虑到第三节的发现——检索占主导地位——    &lt;strong&gt;共享它就是最大的单项节省。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;把读写显式区分开，用笔记的话说，    &lt;strong&gt;最终应该给 agent 超能力。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;XI. 电池候选清单&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#xi-&amp;#30005;&amp;#27744;&amp;#20505;&amp;#36873;&amp;#28165;&amp;#21333;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;笔记最后列了几个可以被原生集成的开源项目，每个都附了「以 Jev 为中心的 harness 会怎么用它」的设计说明。&lt;/p&gt;  &lt;table&gt;    &lt;tr&gt;      &lt;th&gt;项目&lt;/th&gt;      &lt;th&gt;角色&lt;/th&gt;      &lt;th&gt;原生角度&lt;/th&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;headroom&lt;/td&gt;      &lt;td&gt;上下文压缩器&lt;/td&gt;      &lt;td&gt;用分类器检查压缩是否保住了需要的事实&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;rtk&lt;/td&gt;      &lt;td&gt;工具输出压缩&lt;/td&gt;      &lt;td&gt;第一方提示，让模型理解它&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;ast-grep&lt;/td&gt;      &lt;td&gt;结构化搜索&lt;/td&gt;      &lt;td&gt;加载一次手册，生成 N 个查询，按相关性过滤&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;ast-outline&lt;/td&gt;      &lt;td&gt;结构化大纲&lt;/td&gt;      &lt;td&gt;层级调用：选出要检查的子树&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;fastcontext&lt;/td&gt;      &lt;td&gt;仓库探索子 agent&lt;/td&gt;      &lt;td&gt;路由到它，或用结构替代它的搜索&lt;/td&gt;&lt;/tr&gt;    &lt;tr&gt;      &lt;td&gt;fff&lt;/td&gt;      &lt;td&gt;路径与内容搜索&lt;/td&gt;      &lt;td&gt;内存索引、频率排序，长会话里比 ripgrep 快&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;XII. 结论&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#xii-&amp;#32467;&amp;#35770;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;设计笔记提出的论点    &lt;strong&gt;说起来容易、做起来难&lt;/strong&gt;：&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;coding agent 是简单的循环，而杠杆不在循环里。杠杆在于 harness 每一轮放在模型面前的是什么——而今天这个决定是由默认值做的，由一份围绕 KV cache 经济性塑造的、只追加的对话记录做的。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;把缓存拿掉当作思想实验，六个熟悉的行为就变了样：&lt;/p&gt;  &lt;ul&gt;    &lt;li&gt;      &lt;strong&gt;路由失败&lt;/strong&gt;是因为上下文被重新处理，不是因为便宜模型弱&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;工具挤满窗口&lt;/strong&gt;是因为它们必须提前声明&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;Compaction 丢信息&lt;/strong&gt;是因为它在问题已知之前压缩&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;子 agent 稀少&lt;/strong&gt;是因为传状态很难&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;重启&lt;/strong&gt;把好状态和坏状态一起丢掉&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;电池之争&lt;/strong&gt;之所以存在，只是因为每个内置能力永久消耗上下文&lt;/li&gt;&lt;/ul&gt;  &lt;p&gt;所提议的 harness 用一个动作应对全部六项：    &lt;strong&gt;让状态显式且带类型，让 Jev 按查询决定上下文、路由、工具和权限。&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;块在可见性阶梯上被打分。路由按上下文重建定价。工具分级披露。指令附着到条件上。后台任务共享一次检索。&lt;/p&gt;  &lt;p&gt;    &lt;strong&gt;这些都不需要更好的模型。它需要把上下文窗口当作一件被有意组装的东西，而不是一件偶然累积的东西。&lt;/strong&gt;&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;来源说明&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#&amp;#26469;&amp;#28304;&amp;#35828;&amp;#26126;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;p&gt;独立汇编，供研究之用。与 TypeSafe 无隶属关系，亦未获其背书。Jev 依据 TypeSafe 材料描述为一个返回带概率的类型化 choice、score 和 noul 决策的决策模型。&lt;/p&gt;  &lt;p&gt;核心论点、路由算术、六个症状以及所提议的功能，来自 TypeSafe 创始人 Diogo Almeida 提供给汇编者的设计笔记。&lt;/p&gt;  &lt;p&gt;读取与搜索的份额数字（GPT-5.4 轨迹中占工具调用轮次的 56.2%、占主 agent token 的 46.5%）如微软 fastcontext 项目所报告。&lt;/p&gt;  &lt;p&gt;token 份额表是对 CLI coding agent 会话的示意性估计。递归语言模型：A. Zhang, 2025。工具引用：headroom、rtk、ast-grep、ast-outline、fastcontext、fff（GitHub）。模型标价取自笔记所用，可能已经变动。所有图表均为原文所绘。&lt;/p&gt;  &lt;hr&gt;&lt;/hr&gt;  &lt;div&gt;    &lt;h2&gt;关于本翻译&lt;/h2&gt;    &lt;a href="https://github.com/yibie/jev-engineering-zh#&amp;#20851;&amp;#20110;&amp;#26412;&amp;#32763;&amp;#35793;"&gt;&lt;/a&gt;&lt;/div&gt;  &lt;ul&gt;    &lt;li&gt;原文：      &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/assets/original.pdf"&gt;assets/original.pdf&lt;/a&gt;（12 页，WeasyPrint 生成）&lt;/li&gt;    &lt;li&gt;原文纯文本：      &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/assets/original-text.txt"&gt;assets/original-text.txt&lt;/a&gt;&lt;/li&gt;    &lt;li&gt;插图：从原 PDF 的矢量图表按 300 DPI 精确裁切，共 7 张，位于      &lt;code&gt;figures/&lt;/code&gt;&lt;/li&gt;    &lt;li&gt;翻译保留原章节编号与结构；英文原文的排版错误（如小标题的大小写异常      &lt;code&gt;sIX sYMPTOMs OF THE Kv CACHE&lt;/code&gt;）在译文中已按正常书写呈现&lt;/li&gt;    &lt;li&gt;      &lt;strong&gt;上游原始笔记&lt;/strong&gt;：      &lt;a href="https://github.com/yibie/jev-engineering-zh/blob/main/source-notes"&gt;source-notes/&lt;/a&gt;（本文档所依据的设计笔记，及其中文翻译）&lt;/li&gt;&lt;/ul&gt;
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63285-jev-%E5%B7%A5%E7%A8%8B%E5%AD%A6-coding</guid>
      <pubDate>Wed, 23 Sep 2026 16:00:35 CST</pubDate>
    </item>
    <item>
      <title>人工智能软硬件推理优化进行中-Inside the Inference Hardware Revolution Of 2026 - IEEE Spectrum</title>
      <link>https://itindex.net/detail/63284-%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD-%E8%BD%AF%E7%A1%AC-%E6%8E%A8%E7%90%86</link>
      <description>&lt;div&gt;    &lt;p&gt;   &lt;strong&gt;IEEE Spectrum《The AI Inference Revolution Is Here》（Matthew S. Smith，2026年9月15日）核心内容总结&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;文章指出，2020年以来AI的重心在于训练更大的模型，而到2026年，推理（inference）已取代训练成为行业焦点。原因有三：模型真正变得好用、用户激增；如今大量模型是&amp;quot;推理模型&amp;quot;，会通过思维链多次自我提问，高推理强度时输出的文本量可达低强度的约20倍；以及智能体（agentic AI）让推理从实时问答变成全天候自主运行。&lt;/p&gt;  &lt;p&gt;技术上，推理与训练的计算特征不同。由于模型是自回归的，生成每个新词元都要读取全部权重和此前的上下文。回复分为两个阶段：预填充（prefill）并行处理整个提示词、计算注意力，适合GPU；解码（decode）逐词元生成，每预测一个词元都要从内存中读出整个模型，再加上KV缓存的开销。数据搬运所需带宽常常超出硬件能力，导致计算单元空转——有研究发现运行开源大模型的英伟达H100有50%到80%的时间处于闲置状态。&lt;/p&gt;  &lt;p&gt;因此内存成为主战场，各家路线各异：d-Matrix把计算芯片直接堆叠在DRAM之上，将数据传输距离压缩到微米级；Majestic Labs则改进内存接口，用专有铜链路与聚合芯片连接远至约一米外的廉价DRAM，单机架可支持高达128TB内存，远超英伟达GB300 NVL72约20TB的HBM3E；SK海力士等则押注已量产的HBM4。&lt;/p&gt;  &lt;p&gt;巨头选择&amp;quot;分工组合&amp;quot;：英伟达用Rubin GPU处理计算密集的预填充，用片上SRAM丰富的Groq 3 LPU处理内存密集的解码；AWS则将Trainium与Cerebras晶圆级引擎配对，后者把整片晶圆做成单芯片，集成超过4万亿个晶体管和44GB片上SRAM。&lt;/p&gt;  &lt;p&gt;




&lt;/p&gt;  &lt;p&gt;软件侧，量化（如NVFP4、MXFP4）以极小的精度损失换取数倍性能；Tensordyne用对数数系以加法替代乘法，Etched则把Transformer架构直接固化进硅片。文章最后认为，不会有单一赢家——推理硬件将像当年的CPU一样，沿多条路径并行演进。&lt;/p&gt;  &lt;p&gt;   &lt;strong&gt;    &lt;br /&gt;&lt;/strong&gt;&lt;/p&gt;  &lt;p&gt;      &lt;strong&gt;Since about 2020,&lt;/strong&gt;AI has largely focused on training bigger and better models. Large language models (LLMs) ballooned from millions of parameters to trillions. This proved effective: The largest version of OpenAI’s GPT-3, released in 2020, correctly      &lt;a href="https://arxiv.org/pdf/2009.03300" rel="noopener noreferrer" target="_blank"&gt;answered&lt;/a&gt;just 43.9 percent of questions on a popular knowledge-and-reasoning benchmark. Just four years later, GPT-4o      &lt;a href="https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/" rel="noopener noreferrer" target="_blank"&gt;reached&lt;/a&gt;a score of 88.7 percent on the same exam, effectively matching those of human experts.&lt;/p&gt;    &lt;div&gt;&lt;/div&gt;    &lt;p&gt;Advanced AI labs are still training ever larger models, but that training has somewhat receded to the background of the AI conversation. In 2026, inference—the use of trained models to produce code, write essays, or make images of ourselves as elves—has come to the forefront.&lt;/p&gt;    &lt;p&gt;“It’s like training is yesterday’s news,” says      &lt;a href="https://moorinsightsstrategy.com/team/matt-kimball/" target="_blank"&gt;Matt Kimball&lt;/a&gt;, principal data-center analyst at Moor Insights &amp;amp; Strategy. “All that any chief information officer wants to talk about is inference.” Nvidia CEO Jensen Huang, speaking at the company’s GTC 2026 conference, touted this change as the “      &lt;a href="https://www.youtube.com/watch?v=jw_o0xr8MWU" target="_blank"&gt;inflection point of inference&lt;/a&gt;.”&lt;/p&gt;    &lt;p&gt;Part of what’s caused the shift is very simple: LLMs are becoming useful, so people are using them. On top of that, many models on the market today are reasoning models. In response to a user’s query, they run inference not just once but multiple times, reprompting themselves in a process called      &lt;a href="https://arxiv.org/pdf/2201.11903" target="_blank"&gt;        &lt;em&gt;          &lt;em&gt;chain of thought&lt;/em&gt;&lt;/em&gt;&lt;/a&gt;. Reasoning models generate longer outputs, and models with high reasoning effort can produce up to      &lt;a href="https://www.linkedin.com/posts/artificial-analysis_how-many-tokens-do-reasoning-models-use-vs-activity-7318302119206289408-3mt1/" target="_blank"&gt;20 times&lt;/a&gt;as much text as those with low or no effort. Adding even more to the world’s inference workload, the rise of      &lt;a href="https://spectrum.ieee.org/ai-agents" target="_self"&gt;agentic AI&lt;/a&gt;has resulted in inference running not just as a real-time response to a user’s query but also around the clock, working autonomously toward a user-defined goal.&lt;/p&gt;    &lt;p&gt;      &lt;img alt="Close-up of an Annapurna Labs metal processor chip with reflective black surfaces" height="1304" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%201740%201304'%3E%3C/svg%3E" width="1740"&gt;&lt;/img&gt;      &lt;small&gt;Amazon’s Trainium chip was originally designed for AI training. However, Amazon Web Services chose to break up AI inference into two parts, with Trainium running the more computationally complex portion and Cerebras’s wafer-scale engine taking on the more memory-intensive portion.&lt;/small&gt;      &lt;small&gt;Amazon&lt;/small&gt;&lt;/p&gt;    &lt;p&gt;The resulting explosion in inference demand has led to unexpected alliances among tech giants.      &lt;a href="https://openai.com/index/cerebras-partnership/" target="_blank"&gt;OpenAI&lt;/a&gt;and      &lt;a href="https://www.reuters.com/business/retail-consumer/cerebras-systems-amazon-strike-deal-offer-cerebras-ai-chips-amazons-cloud-2026-03-13/" target="_blank"&gt;Amazon&lt;/a&gt;have deployed chips the size of a      &lt;a href="https://spectrum.ieee.org/cerebrass-giant-chip-will-smash-deep-learnings-speed-barrier" target="_self"&gt;dinner plate&lt;/a&gt;designed by      &lt;a href="https://www.cerebras.ai/" target="_blank"&gt;Cerebras&lt;/a&gt;, despite Amazon having its own      &lt;a href="https://aws.amazon.com/ai/machine-learning/trainium/" target="_blank"&gt;Trainium&lt;/a&gt;chips. Nvidia      &lt;a href="https://www.cnbc.com/2025/12/24/nvidia-buying-ai-chip-startup-groq-for-about-20-billion-biggest-deal.html" target="_blank"&gt;bought&lt;/a&gt;key talent and intellectual property from AI-inference startup      &lt;a href="https://groq.com/" target="_blank"&gt;Groq&lt;/a&gt;in a controversial deal worth US $20 billion. And      &lt;a href="https://www.anthropic.com/" target="_blank"&gt;Anthropic&lt;/a&gt;is      &lt;a href="https://x.ai/news/anthropic-compute-partnership" target="_blank"&gt;paying&lt;/a&gt;LLM competitor      &lt;a href="https://x.ai/" target="_blank"&gt;SpaceXAI&lt;/a&gt;over a billion dollars per month to lease spare compute.&lt;/p&gt;    &lt;p&gt;Although they might seem similar, AI training and AI inference are computationally different. These big moves from tech giants signal that in order to support the inference demand, we’re going to need a very different mix of hardware than experts may have expected even a couple of years ago.&lt;/p&gt;    &lt;h2&gt;How does AI inference differ from AI training?&lt;/h2&gt;    &lt;p&gt;An untrained LLM is like a jumble of Scrabble tiles on a table. Instead of single letters, though, the tiles show fragments of words, called tokens. Everything you’d need to write almost anything is present, but nothing makes sense.&lt;/p&gt;    &lt;p&gt;Training a model organizes this jumble using a guessing game played at scale. The model is shown real text with the next token hidden and asked to predict what comes next. After each guess, the correct token is revealed and then compared to the prediction, and the difference is used to calculate the model’s accuracy. The game is played not with a single sentence but over billions of passages.&lt;/p&gt;    &lt;p&gt;&lt;/p&gt;    &lt;div&gt;      &lt;p&gt;While a real game of Scrabble can be played over a bag of chips and a few drinks, AI training is computationally intense. The model updates its parameters through        &lt;a href="https://spectrum.ieee.org/what-is-deep-learning/backpropagation" target="_self"&gt;backpropagation&lt;/a&gt;, a process that repeatedly calculates how each of a model’s billions or trillions of parameters should shift to make the next prediction better. This is why tech giants are        &lt;a href="https://spectrum.ieee.org/5gw-data-center" target="_self"&gt;building&lt;/a&gt;larger data centers than ever before.&lt;/p&gt;      &lt;p&gt;Eventually the model’s creator decides further training isn’t worth the cost, and the guessing game stops. Backpropagation ends, the parameters are frozen, and the LLM becomes a pretrained model. Fine-tuning—a short training run on smaller, more specialized data—adds final tweaks, and the model is deployed.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;img alt="Close-up of a gold computer chip with rainbow-colored circuitry on black background" height="1499" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%202000%201499'%3E%3C/svg%3E" width="2000"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Nvidia’s Groq 3 language-processing unit minimizes data movement by placing on-chip SRAM memory and computational blocks in the order they are needed on-chip.&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Nvidia&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;Next comes inference. This is the process of using the deployed model, which, now that it’s been trained, has learned to spit out Scrabble tiles—tokens—in a sensible order.&lt;/p&gt;      &lt;p&gt;You might think that AI inference is less computationally demanding because the backpropagation calculations used to update parameters are eliminated. But        &lt;a href="https://www.linkedin.com/in/sudeep-bhoja-070a111/" target="_blank"&gt;Sudeep Bhoja&lt;/a&gt;, founder and CTO of the inference-hardware company        &lt;a href="https://www.d-matrix.ai/" rel="noopener noreferrer" target="_blank"&gt;d-Matrix&lt;/a&gt;, explains that inference adds new challenges.&lt;/p&gt;      &lt;p&gt;The models are “autoregressive” in nature. That is, the next output depends on the previous one. “So to generate the next token, you have to read all of the weights and all of the [context] from the previous token,” explains Bhoja. The context includes all of your prompts, all of the LLM’s replies, and all of the files you upload. It’s a lot of data and a lot of processing.&lt;/p&gt;      &lt;p&gt;An LLM generates its reply in two phases: prefill and decode. Prefill is the model reading a prompt. It processes every token at once, computing how each token relates to all the others. This operation is called        &lt;a href="https://en.wikipedia.org/wiki/Attention_(machine_learning)" rel="noopener noreferrer" target="_blank"&gt;attention&lt;/a&gt;, and it’s a defining characteristic of the transformer architecture behind modern LLMs. It allows them to respond to a word in its sentence, paragraph, and larger context rather than on its own. Think of it like arranging Scrabble tiles before you place them in a game. Many players move tiles around to imagine how they connect. Self-attention plays a similar role, though instead of moving physical tiles, each token sends a query to the others and receives a score indicating the token’s relevance.&lt;/p&gt;      &lt;p&gt;These queries result in two types of vectors: the keys and values. They are typically placed in a store called the KV cache. This isn’t strictly required, as a model could instead recompute these vectors with each new token it generates. But nearly all LLMs use a KV cache to reduce how much computing they do. The KV cache is stored in memory and becomes a scratchpad to which the LLM can return to understand a conversation, and though it starts small, it can swell to dozens of gigabytes.&lt;/p&gt;      &lt;p&gt;Prefill is a problem that can be easily divided up and worked on in parallel. This is why GPUs became the dominant AI accelerator as LLMs surged in popularity. Graphics rasterization (computing the color of every pixel on a screen) is also massively parallel, so GPU architectures were a natural fit.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;img alt="Gloved hands holding a large golden computer processor wafer" height="1500" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%202000%201500'%3E%3C/svg%3E" width="2000"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Cerebras’s wafer-scale engine chips maximize memory bandwidth by keeping everything—both memory and computational units—side by side on the dinner-plate-size chips.          &lt;br /&gt;&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Cerebras&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;Next comes decode. Here, the model generates its reply one token at a time. At each step it takes the most recent token, weighs it against everything in the KV cache, uses that information to predict the next token, and adds the new token’s key and value to the cache. Then it repeats in sequence, token by token.&lt;/p&gt;      &lt;p&gt;This is where the autoregressive nature of the model works against inference speed. Predicting each token requires reading the entire model from memory, and that model consists of possibly tens to hundreds of gigabytes of parameters (the numbers representing what the model learned in training). Crucially, this is in addition to the memory required to store the KV cache.&lt;/p&gt;      &lt;p&gt;As a result, the movement of all this data through memory often requires more bandwidth than inference hardware has available. So at least some of the computing parts of a GPU sit idle as it waits for data. Researchers        &lt;a href="https://arxiv.org/pdf/2503.08311" target="_blank"&gt;found&lt;/a&gt;that Nvidia H100 GPUs running open-source LLMs sit idle 50 to 80 percent of the time.&lt;/p&gt;      &lt;h2&gt;Memory’s role in inferencing&lt;/h2&gt;      &lt;p&gt;        &lt;a href="https://www.linkedin.com/in/rabii/" target="_blank"&gt;Shahriar “Sha” Rabii&lt;/a&gt;, former head of silicon engineering at Meta and cofounder of the AI startup        &lt;a href="https://majestic-labs.ai/" target="_blank"&gt;Majestic Labs&lt;/a&gt;, says idled processors are why many companies that are trying to improve AI-inference performance are laser-focused on memory. “With the GPU-based approach, you end up greatly over-provisioning compute and starved on memory. That’s driving the big [memory] scale out,” he says.&lt;/p&gt;      &lt;p&gt;Bhoja’s d-Matrix and Rabii’s Majestic Labs both focus on this memory bottleneck. However, their companies imagine different solutions.&lt;/p&gt;      &lt;p&gt;d-Matrix’s second-generation AI accelerator,        &lt;a href="https://www.d-matrix.ai/announcements/d-matrix-and-alchip-announce-collaboration-on-worlds-first-3d-dram-solution-to-supercharge-ai-inference/" rel="noopener noreferrer" target="_blank"&gt;Raptor&lt;/a&gt;, aims to improve inference performance by minimizing the distance between compute and memory. The GPUs in most current AI-inference deployments do this by placing high-bandwidth memory (HBM) around the perimeter of the GPU. Each HBM is a stack of DRAM dies linked together and connected to a superfast interface to the GPU. This is great for training, but for inference, the amount of memory you can stack this way and the bandwidth it can provide leave something to be desired.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;h3&gt;d-Matrix’s stacked-die architecture&lt;/h3&gt;      &lt;img alt="Diagram of stacked logic and DRAM chips connected by solder bumps on a substrate" height="361" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%201200%20361'%3E%3C/svg%3E" width="1200"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Memory bandwidth—how quickly data can be read from memory to logic—is a major bottleneck in AI inference. The startup d-Matrix is increasing that bandwidth by stacking the logic die directly on top of the memory, in this case DRAM. This allows for lots of extremely short interconnects.&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Chris Philpot&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;        &lt;br /&gt;&lt;/p&gt;      &lt;p&gt;d-Matrix’s Raptor removes that bottleneck by stacking an AI accelerator on a DRAM die. Instead of stacking memory, d-Matrix stacks memory and compute. Bhoja says this reduces the distance that data must travel to “micrometers instead of millimeters.” Like building a skyscraper, going vertical makes it possible to do more inside the same physical footprint.&lt;/p&gt;      &lt;p&gt;Majestic takes the opposite approach. Instead of trying to minimize the length that data must travel between compute and memory, the company is focused on improving the memory interface to accommodate longer wire traces while keeping bandwidth high. Longer wires allow Majestic to connect memory stacks that aren’t directly next to the GPU, removing the space limitation of HBM.&lt;/p&gt;      &lt;p&gt;“A memory interface has a very short physical distance it can operate over. In the case of HBM, it’s up to 2 or 3 millimeters. You have this shoreline around the periphery, which is the only place where you can put HBM,” says Rabii.&lt;/p&gt;      &lt;p&gt;Majestic        &lt;a href="https://www.techradar.com/pro/startup-swaps-costly-ai-gpus-for-arm-cores-and-up-to-128tb-of-cheap-lpddr6-ram-instead-of-expensive-hbm-to-smash-through-the-memory-wall" rel="noopener noreferrer" target="_blank"&gt;claims&lt;/a&gt;its memory interface can transmit bits as far as about a meter. That’s achieved with a proprietary copper link and a memory-aggregator chip that coordinates data. “The aggregator is the endpoint for the high-speed interface and a way to fan out to many, many commodity DRAM chips,” says Rabii. Because of this, Majestic can support up to 128 terabytes of DRAM memory in a single server rack—a significant increase over Nvidia’s        &lt;a href="https://www.nvidia.com/en-us/data-center/gb300-nvl72/" rel="noopener noreferrer" target="_blank"&gt;GB300 NVL72 rack&lt;/a&gt;, which has about        &lt;a href="https://resources.nvidia.com/en-us-blackwell-architecture/blackwell-ultra-datasheet?ncid=no-ncid" rel="noopener noreferrer" target="_blank"&gt;20 TB of HBM3E&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;h3&gt;Majestic Labs’ memory-aggregation architecture&lt;/h3&gt;      &lt;img alt="Diagram of memory aggregator chiplet linking server GPUs/CPUs to shared DRAM pool" height="604" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%201200%20604'%3E%3C/svg%3E" width="1200"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Majestic Labs plans to satisfy AI’s memory appetite via a proprietary interconnect and a memory-aggregator chip, allowing a single rack access to up to 128 terabytes of cheap DRAM memory.&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Chris Philpot&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;d-Matrix and Majestic have one thing in common: Instead of HBM, they both use off-the-shelf DRAM. This is the most common type of computer memory in the world; it’s in everything from smartphones to cars. Memory analyst        &lt;a href="https://thememoryguy.com/" target="_blank"&gt;Jim Handy&lt;/a&gt;says HBM costs two to three times as much as DRAM. d-Matrix and Majestic chose DRAM in part because of this price advantage. However, the proponents of HBM, which include memory giants like        &lt;a href="https://www.samsung.com/us/" target="_blank"&gt;Samsung&lt;/a&gt;and        &lt;a href="https://www.skhynix.com/" target="_blank"&gt;SK Hynix&lt;/a&gt;, aren’t sitting idle.&lt;/p&gt;      &lt;p&gt;HBM4, the latest version of HBM memory, is now in production and will be used by        &lt;a href="https://spectrum.ieee.org/nvidia-rubin-networking" target="_blank"&gt;Nvidia’s Vera Rubin GPU&lt;/a&gt;, which is expected to ship in the second half of 2026.        &lt;a href="https://www.linkedin.com/in/hoshikk/" target="_blank"&gt;Hoshik Kim&lt;/a&gt;, head of memory-systems research at        &lt;a href="https://www.skhynix.com/" rel="noopener noreferrer" target="_blank"&gt;SK Hynix&lt;/a&gt;, says HBM4 “will decisively break the memory bottlenecks constraining AI inference today” by doubling HBM’s maximum memory bandwidth and increasing the amount of HBM memory per stack.&lt;/p&gt;      &lt;h2&gt;Combining chips for faster inference&lt;/h2&gt;      &lt;p&gt;The big players—Nvidia and Amazon—are going for an all-chips-on-deck approach. Nvidia’s GPUs and Amazon’s Trainium training accelerators are still great for part of the inference workload: the prefill stage, where all the context keys and values are calculated. But to accelerate decode, the part where new tokens are generated, they are looking to new, memory-centric architectures from smaller players.&lt;/p&gt;      &lt;p&gt;In Nvidia’s case, the smaller player was Groq (not to be confused with Grok, the family of LLMs trained by SpaceXAI). Nvidia purchased intellectual property and hired talent from Groq at the end of 2025, and just three months later at the Nvidia’s GTC 2026 conference, Jensen Huang        &lt;a href="https://spectrum.ieee.org/nvidia-groq-3" target="_self"&gt;unveiled&lt;/a&gt;the Nvidia Groq 3 language-processing unit (        &lt;a href="https://www.nvidia.com/en-us/data-center/lpx/" rel="noopener noreferrer" target="_blank"&gt;LPU&lt;/a&gt;). Groq’s architecture relies on memory—in its case, SRAM—built directly into the chip’s architecture.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;Unless you’re a chip architect, or a        &lt;a href="https://www.pcworld.com/article/2634140/why-i-care-about-cpu-cache-as-a-pc-gamer-the-obscure-spec-explained.html" target="_blank"&gt;hardcore PC gamer,&lt;/a&gt;you probably never give SRAM a thought. SRAM has the benefit of being tightly integrated into a compute chip’s architecture—it’s on the same piece of silicon as the processor—and has the drawback of being less dense and more expensive than DRAM. Most chips include only a few dozen megabytes of SRAM. AI inference, however, has ignited new interest in SRAM as a means of bringing the model weights stored in memory closer to compute.&lt;/p&gt;      &lt;p&gt;        &lt;a href="https://www.linkedin.com/in/ian-buck-19201315/" target="_blank"&gt;Ian Buck&lt;/a&gt;, vice-president and general manager of hyperscale and high-performance computing at        &lt;a href="https://blogs.nvidia.com/blog/author/ian-buck/" target="_blank"&gt;Nvidia&lt;/a&gt;, says the LPU has a much different set of priorities than the company’s GPUs. The LPU has far less raw computing power than a standard GPU, but it gains 500 megabytes of on-die SRAM connected directly to its floating-point math units. “The benefit is the memory bandwidth. The LPU has seven times the memory bandwidth of the GPU,” he says.&lt;/p&gt;      &lt;p&gt;Between the Rubin GPU and the Groq LPU, prefill and decode can both be accelerated to get the best of both worlds, the theory goes. “We do all the attention math and context processing on the Vera Rubin [GPU] rack,” explains Buck. “For all the expert calculations…the matrix multiplications, we do that part on the LPU.” The company packs 256 LPUs into the Groq 3 LPX, a system the size of a data-center rack.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;h3&gt;Nvidia’s two-chip approach to inference&lt;/h3&gt;      &lt;img alt="Diagram comparing Nvidia Rubin GPU and Groq 3 LPU chip layouts with labeled blocks" height="666" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%201447%20666'%3E%3C/svg%3E" width="1447"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Nvidia also plans to split the inference workload across two chips. The company’s newest Rubin GPUs will tackle the compute-intensive prefill phase, while the Groq 3 language-processing unit (LPU), with lots of on-chip SRAM, will handle the memory-intensive decode phase.&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Chris Philpot&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;Amazon Web Services (AWS), for its part, struck a        &lt;a href="https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference" target="_blank"&gt;deal&lt;/a&gt;with        &lt;a href="https://spectrum.ieee.org/tag/cerebras"&gt;Cerebras&lt;/a&gt;, to pair the Trainium accelerator with        &lt;a href="https://www.cerebras.ai/chip" target="_blank"&gt;Cerebras’s Wafer-Scale Engine 3 (WSE-3)&lt;/a&gt;. Cerebras takes a similar approach to Groq, though at a much larger scale. WSE-3 turns an entire silicon wafer into a single chip that contains over 4 trillion transistors. The design doesn’t connect to external memory but instead etches 44 gigabytes of SRAM into each wafer. “We store the [model] weights on the SRAM,” says        &lt;a href="https://www.linkedin.com/in/james-wang-5166575/" target="_blank"&gt;James Wang&lt;/a&gt;, formerly director of product marketing at Cerebras who has since moved to SpaceXAI. “So that’s easily 40 to up to 80 billion parameters that we can support on one chip.”&lt;/p&gt;      &lt;p&gt;Amazon plans to use AWS Trainium chips for prefill, and Cerebras for decode. But Cerebras’s chips can also go it alone in inference. WSE-3 was        &lt;a href="https://openai.com/index/introducing-gpt-5-3-codex-spark/" target="_blank"&gt;deployed by OpenAI to power GPT-5.3-Codex-Spark&lt;/a&gt;, a variant of the company’s coding mode, outputting over 1,000 tokens per second. For comparison, OpenAI’s standard GPT-5.4 deployment outputs 50 to 125 tokens per second.&lt;/p&gt;&lt;/div&gt;    &lt;div&gt;      &lt;h3&gt;Amazon Web Services’ two-chip inference strategy &lt;/h3&gt;      &lt;img alt="A schematic of Amazon's Trainium chip on the left, with SRAM memory block and logic blocks plus high-bandwidth memory. Schematic of Cerebras's wafer-scale engine on right, with small SRAM memory and logic block in a checkerboard pattern." height="478" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%201199%20478'%3E%3C/svg%3E" width="1199"&gt;&lt;/img&gt;      &lt;small&gt;        &lt;p&gt;Amazon Web Services combined their Trainium chips with Cerebras’s dinner-plate-sized wafer-scale engine (WSE) to tackle different parts of AI inference. Trainium chips handle the computationally intensive prefill phase, while the WSE, with interleaved on-chip SRAM memory, handles the memory-bandwidth-limited decode phase.          &lt;br /&gt;&lt;/p&gt;&lt;/small&gt;      &lt;small&gt;        &lt;p&gt;Chris Philpot&lt;/p&gt;&lt;/small&gt;&lt;/div&gt;    &lt;div&gt;      &lt;p&gt;Cerebras can also tackle prefill without moving the workload to different specialized chips. For this, it networks together multiple WSE-3 chips to form a single pool of memory. Cerebras has demonstrated it can serve models with up to 1T parameters, such as        &lt;a href="https://spectrum.ieee.org/tag/moonshot"&gt;Moonshot&lt;/a&gt;AI’s Kimi 2.6, though Wang says “the architecture has no innate limitation in terms of how many parameters it will do.”&lt;/p&gt;      &lt;p&gt;Despite these differences in strategy, Nvidia and AWS seem to agree that the future of AI inference will be solved by a systems approach that pools different kinds of chips together to tackle the largest LLMs. Or, as Buck says: “To do modern AI inference, you need all the chips.”&lt;/p&gt;      &lt;h2&gt;Learning to do more with less (bits)&lt;/h2&gt;      &lt;p&gt;Nvidia became the world’s most valuable tech company because it designed the world’s most desired GPUs. But not all of the attention is focused on improving AI-inference hardware. AI researchers are also learning how to optimize LLM software and hardware in tandem to make the best use of the memory and compute components.&lt;/p&gt;      &lt;p&gt;Most computers store numbers in a 32-bit or 64-bit format. These determine how many bits are available to represent a single number. If too few bits are available, the number can’t be stored without losing information. The quality of an LLM benefits from more-precise number formats, but this creates a problem for inference performance. More-precise numbers aren’t free. The bits that describe them take up more space in memory and require more silicon and energy to compute.&lt;/p&gt;      &lt;p&gt;        &lt;a href="https://www.linkedin.com/in/gillesbackhus/?originalSubdomain=de" target="_blank"&gt;Gilles Backhus&lt;/a&gt;, cofounder of the AI-accelerator company        &lt;a href="https://www.tensordyne.ai/" target="_blank"&gt;Tensordyne&lt;/a&gt;, says this creates a tension between model size and number precision. “Would you prefer a model that is size        &lt;em&gt;x&lt;/em&gt;but runs in 8-bit, or would you prefer a model that is twice the size but runs in 4-bit?” The size of each model will be roughly the same in terms of memory and compute, “but the 4-bit approach gives you twice as many synapses, if you will. And people are figuring out that [the 4-bit approach] is worth it.”&lt;/p&gt;&lt;/div&gt;    &lt;p&gt;&lt;/p&gt;    &lt;p&gt;The process of converting an LLM from a more-precise number format to a less-precise format is called      &lt;a href="https://spectrum.ieee.org/1-bit-llm" target="_self"&gt;quantization&lt;/a&gt;, and it’s been in use for several years. However, researchers are finding new ways to quantize models down while retaining a large majority of the model’s quality.&lt;/p&gt;    &lt;p&gt;Nvidia recently created a new 4-bit number format,      &lt;a href="https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/" target="_blank"&gt;NVFP4&lt;/a&gt;, for this purpose.      &lt;a href="https://www.amd.com/en.html" target="_blank"&gt;AMD&lt;/a&gt;,      &lt;a href="https://www.intel.com/content/www/us/en/homepage.html" target="_blank"&gt;Intel&lt;/a&gt;, and      &lt;a href="https://www.qualcomm.com/" target="_blank"&gt;Qualcomm&lt;/a&gt;have instead rallied around a competing 4-bit number format called      &lt;a href="https://huggingface.co/blog/RakshitAralimatti/learn-ai-with-me" target="_blank"&gt;MXFP4&lt;/a&gt;that Nvidia also contributed to developing. “It’s the black art of AI,” says Buck, of Nvidia. When Nvidia quantized DeepSeek-R1 from FP8 to NVFP4, scores on seven major benchmarks degraded by less than one percent while      &lt;a href="https://developer.nvidia.com/blog/3-ways-nvfp4-accelerates-ai-training-and-inference/" target="_blank"&gt;performance improved by three times&lt;/a&gt;, the company says.&lt;/p&gt;    &lt;p&gt;Quantization is likely just the tip of the spear, as AI researchers and startups are investigating a diversity of opportunities for optimization, some of which could dramatically change the silicon found in AI-inference hardware.&lt;/p&gt;    &lt;p&gt;      &lt;img alt="TENSORDYNE TDN AIP chip with central green processor cores on black board" height="1500" src="data:image/svg+xml,%3Csvg%20xmlns='http://www.w3.org/2000/svg'%20viewBox='0%200%202000%201500'%3E%3C/svg%3E" width="2000"&gt;&lt;/img&gt;      &lt;small&gt;Tensordyne’s unique approach to AI inference combines a logarithmic number format with bespoke hardware in the company’s Napier chip.&lt;/small&gt;      &lt;small&gt;Tensordyne&lt;/small&gt;&lt;/p&gt;    &lt;p&gt;Tensordyne is expected to      &lt;a href="https://spectrum.ieee.org/tensordyne-inference-claim" target="_self"&gt;accelerate&lt;/a&gt;AI inference with a logarithmic number system that leans on a property of logarithms: The log of A times B equals the log of A plus the log of B. So, storing numbers as their exponents lets the chip add where it would otherwise multiply. That matters in silicon because multiplier circuits draw more power and use more die area than adders do. Tensordyne says its rack-scale hardware, called Napier, can produce up to 1,300 tokens per second per user, and can do so while using less than a      &lt;a href="https://www.tensordyne.ai/stories/tensordyne-announces-breakthrough-inference-system-to-end-ais-speed-vs-cost-trade-off" target="_blank"&gt;tenth&lt;/a&gt;as much power as comparable Nvidia hardware.&lt;/p&gt;    &lt;p&gt;      &lt;a href="https://www.etched.com/" target="_blank"&gt;Etched&lt;/a&gt;, a startup based in San Jose, Calif., is even designing AI accelerators that translate the transformer architecture used by LLMs directly into silicon. Rather than building general-purpose GPUs, the company is wiring up the connections needed for efficient transformer calculations into its chip, making the chip much less flexible but more efficient for the tasks most performed by current LLMs. The company says its first AI accelerator,      &lt;a href="https://www.spheron.network/blog/etched-ai-sohu-vs-nvidia-transformer-asic-inference/" target="_blank"&gt;Sohu&lt;/a&gt;, can run Meta’s      &lt;a href="https://spectrum.ieee.org/tag/llama"&gt;Llama&lt;/a&gt;70B model at a stunning 500,000 tokens per second, though this approach also means it won’t be able to run LLMs that move away from a typical transformer architecture.&lt;/p&gt;    &lt;p&gt;Whether these ideas will prove fruitful remains to be seen. Etched just      &lt;a href="https://www.etched.com/progress/from-zero-to-one" target="_blank"&gt;shipped&lt;/a&gt;their first rack in August. Tensordyne believes its first hardware will be available in 2027. Even so, these startups show how the demand for inference performance is fueling unconventional ideas.&lt;/p&gt;    &lt;h2&gt;Inference is everyone’s game&lt;/h2&gt;    &lt;p&gt;The sheer variety of approaches to AI-inference acceleration—stacking compute on memory, extending interfaces from millimeters to meters, using an entire silicon wafer for SRAM, squeezing models into 4 bits—raises a question: Which is going to win, and which is going to lose?&lt;/p&gt;    &lt;p&gt;But that’s likely not the right question, experts say. The demand for AI is currently insatiable, and while fears of an AI bubble stalk the industry, it has yet to hamper growth.&lt;/p&gt;    &lt;p&gt;On the contrary, Kimball of Moor Insights &amp;amp; Strategy thinks inference could drive intense demand for AI hardware in the long term, because it’s not obvious where that demand will end. “You could add a million agents into your organization,” he says. “These things work 24 hours a day; they don’t go home at five at night like we do.”&lt;/p&gt;    &lt;p&gt;If AI inference remains as desirable as Kimball expects, the evolution is likely to follow the same trajectory as the CPU. The CPU didn’t improve along a single axis but instead across      &lt;a href="https://spectrum.ieee.org/intel-i860" target="_self"&gt;multiple fronts&lt;/a&gt;simultaneously. Once transistor scaling slowed, chip and system architecture innovations of all kinds proliferated. The list of individual innovations that led to today’s ubiquitous, powerful personal compute could fill dozens of books.&lt;/p&gt;    &lt;p&gt;A few decades from now, the history of AI inference innovation will show similar depth.&lt;/p&gt;    &lt;div&gt;&lt;/div&gt;    &lt;div&gt;      &lt;div&gt;From Your Site Articles&lt;/div&gt;      &lt;ul&gt;        &lt;li&gt;          &lt;a href="https://spectrum.ieee.org/nvidia-groq-3" rel="noopener noreferrer" target="_blank"&gt;With Nvidia Groq 3, the Era of AI Inference Is (Probably) Here ›&lt;/a&gt;&lt;/li&gt;        &lt;li&gt;          &lt;a href="https://spectrum.ieee.org/new-inference-chips" rel="noopener noreferrer" target="_blank"&gt;AI Inference Competition Heats Up ›&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;      &lt;div&gt;Related Articles Around the Web&lt;/div&gt;      &lt;ul&gt;        &lt;li&gt;          &lt;a href="https://www.siliconflow.com/articles/the-fastest-ai-inference-engine" rel="noopener noreferrer" target="_blank"&gt;Ultimate Guide – The Best and Fastest AI Inference Engines of 2026 - SiliconFlow ›&lt;/a&gt;&lt;/li&gt;        &lt;li&gt;          &lt;a href="https://gradientflow.substack.com/p/llm-inference-hardware-emerging-from" rel="noopener noreferrer" target="_blank"&gt;LLM Inference Hardware: Emerging from Nvidia&amp;apos;s Shadow ›&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/div&gt;&lt;/div&gt;
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63284-%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD-%E8%BD%AF%E7%A1%AC-%E6%8E%A8%E7%90%86</guid>
      <pubDate>Wed, 16 Sep 2026 09:30:27 CST</pubDate>
    </item>
    <item>
      <title>养狗有助于降低老人患认知症风险</title>
      <link>https://itindex.net/detail/63283-%E8%80%81%E4%BA%BA-%E8%AE%A4%E7%9F%A5-%E9%A3%8E%E9%99%A9</link>
      <description>日本国立环境研究所等机构从 2016 年起，历时 7 年半对约 1.1 万名老年人开展了调查。他们在学术期刊上发表了研究成果。养狗的老年人因认知症需要接受护理的风险比从未养狗的人群低 48%。研究认为，遛狗带来的身体活动以及社交往来起到了积极作用。曾经养过狗的人患认知症的风险也低于从未养过狗的人群。虽然该差异在统计学上并不显著，但推测养狗时期建立的人际联系等因素可能带来了积极影响。研究还表明，养狗能拉动经济。若养狗人群增加，宠物食品、宠物保险、宠物寄养等相关商品与服务的需求预计随之上涨。
 &lt;p&gt;&lt;/p&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63283-%E8%80%81%E4%BA%BA-%E8%AE%A4%E7%9F%A5-%E9%A3%8E%E9%99%A9</guid>
      <pubDate>Mon, 14 Sep 2026 18:49:35 CST</pubDate>
    </item>
    <item>
      <title>Cloudflare 免费服务使用指南</title>
      <link>https://itindex.net/detail/63282-cloudflare-%E5%85%8D%E8%B4%B9-%E6%9C%8D%E5%8A%A1</link>
      <description>&lt;p&gt;Cloudflare 的免费套餐因其慷慨的额度，在开发者社区中常被戏称为“赛博大善人”，它提供的服务远不止基础的 CDN 加速。&lt;/p&gt;
 &lt;h2&gt;Cloudflare 提供的免费服务&lt;/h2&gt;
 &lt;p&gt;  &lt;img alt="" height="386" src="https://www.biaodianfu.com/wp-content/uploads/2026/09/Cloudflare.png" width="785"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;Cloudflare 免费层能覆盖的内容，按「值得用」排序：&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;    &lt;strong&gt;优先级&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;    &lt;strong&gt;用途&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;    &lt;strong&gt;免费层是否够用&lt;/strong&gt;&lt;/td&gt;
   &lt;td&gt;    &lt;strong&gt;关键限制&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1&lt;/td&gt;
   &lt;td&gt;域名 DNS 托管 + 免费 HTTPS&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;无实质限制&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2&lt;/td&gt;
   &lt;td&gt;CDN 缓存 + 基础 DDoS 防护&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;无流量计费&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;3&lt;/td&gt;
   &lt;td&gt;静态站托管（Pages）&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;500 次构建/月&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;4&lt;/td&gt;
   &lt;td&gt;内网服务暴露（Tunnel + Access）&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;50 用户上限&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;5&lt;/td&gt;
   &lt;td&gt;域名邮箱转发（Email Routing）&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;只能转发，不能发信&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;6&lt;/td&gt;
   &lt;td&gt;反爬 / 验证码（Turnstile）&lt;/td&gt;
   &lt;td&gt;完全够用&lt;/td&gt;
   &lt;td&gt;20 个 widget&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;7&lt;/td&gt;
   &lt;td&gt;边缘 API（Workers）&lt;/td&gt;
   &lt;td&gt;小流量够用&lt;/td&gt;
   &lt;td&gt;10 万请求/天、10ms CPU&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;8&lt;/td&gt;
   &lt;td&gt;对象存储（R2）&lt;/td&gt;
   &lt;td&gt;小文件量够用&lt;/td&gt;
   &lt;td&gt;10 GB 存储、零出口费&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;9&lt;/td&gt;
   &lt;td&gt;数据库（D1 + KV）&lt;/td&gt;
   &lt;td&gt;原型够用&lt;/td&gt;
   &lt;td&gt;500 万行读/天&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;10&lt;/td&gt;
   &lt;td&gt;AI 推理（Workers AI）&lt;/td&gt;
   &lt;td&gt;只够试玩&lt;/td&gt;
   &lt;td&gt;1 万 Neurons/天&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;一句话判断标准：  &lt;strong&gt;只分发不计算，免费层几乎无限；一旦要算、要存、要实时，就会撞墙。&lt;/strong&gt;&lt;/p&gt;
 &lt;h3&gt;为什么值得用：免费层的底层逻辑&lt;/h3&gt;
 &lt;p&gt;理解 Cloudflare 的定价哲学，比背额度表有用得多。&lt;/p&gt;
 &lt;p&gt;它把成本结构拆成两块：&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;成本类型&lt;/td&gt;
   &lt;td&gt;定价态度&lt;/td&gt;
   &lt;td&gt;原因&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;带宽 / 分发 / 连接&lt;/td&gt;
   &lt;td&gt;几乎免费送&lt;/td&gt;
   &lt;td&gt;边缘节点带宽边际成本极低，且流量能反哺网络规模&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CPU 时间 / 存储 / 写入&lt;/td&gt;
   &lt;td&gt;严格按量计费&lt;/td&gt;
   &lt;td&gt;这是真实的可变成本&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;这解释了为什么 Pages 敢写「无限带宽、无限请求」，而 Workers 只给 10 毫秒 CPU——带宽对它来说接近免费，CPU 不是。&lt;/p&gt;
 &lt;p&gt;对应的三条推论：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;静态内容越重，越占便宜。图片站、文档站、博客、下载站放在这里，成本趋近于零。&lt;/li&gt;
  &lt;li&gt;动态逻辑越重，越容易付费。视频转码、爬虫、长任务、复杂数据库查询，很快就会越过免费线。&lt;/li&gt;
  &lt;li&gt;免费层的天花板由并发而非流量决定。1000 请求/分钟的突发限制，比每天 10 万次的总量更早撞上。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;免费额度是硬限制。超出后请求直接失败并返回错误码（Workers 是 1015 / 1027），不是排队也不是降速。这意味着它适合个人站、演示、原型，不适合当作有 SLA 承诺的生产依赖。&lt;/p&gt;
 &lt;h3&gt;免费产品地图&lt;/h3&gt;
 &lt;p&gt;Cloudflare 免费层大致分五层。理解分层，就知道该在哪个位置放什么。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;域名层：DNS + TLS + CDN + DDoS&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;能力&lt;/td&gt;
   &lt;td&gt;免费层说明&lt;/td&gt;
   &lt;td&gt;核实状态&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;权威 DNS&lt;/td&gt;
   &lt;td&gt;Anycast 全网解析，支持 API 与常见记录类型&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;免费 TLS 证书&lt;/td&gt;
   &lt;td&gt;自动签发、自动续期，覆盖根域与泛域名（泛域名需付费）&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CDN 缓存&lt;/td&gt;
   &lt;td&gt;静态资源缓存、缓存规则、Cache Rules 基本能力&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;DDoS 防护&lt;/td&gt;
   &lt;td&gt;L3/L4 不限量清洗，L7 基础防护&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;WAF&lt;/td&gt;
   &lt;td&gt;自定义规则、免费托管规则集、IP 访问规则、UA 封禁&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;速率限制&lt;/td&gt;
   &lt;td&gt;Free 计划 1 条规则&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;这一层是免费层的入口，也是性价比最高的一层。绝大多数人只需要做一件事：把域名 nameserver 改成 Cloudflare 的。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;托管层：Pages + Workers&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;项目&lt;/td&gt;
   &lt;td&gt;Workers Free&lt;/td&gt;
   &lt;td&gt;Pages Free&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;请求&lt;/td&gt;
   &lt;td&gt;100,000 次/天（账号内所有脚本合计）&lt;/td&gt;
   &lt;td&gt;静态请求不限量&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;CPU 时间&lt;/td&gt;
   &lt;td&gt;10 ms / 次调用&lt;/td&gt;
   &lt;td&gt;静态资源无 CPU 消耗&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;突发限制&lt;/td&gt;
   &lt;td&gt;1,000 请求/分钟&lt;/td&gt;
   &lt;td&gt;静态资源不适用&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;脚本 / 项目数&lt;/td&gt;
   &lt;td&gt;最多 100 个 Worker 脚本&lt;/td&gt;
   &lt;td&gt;每账号 100 个项目&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;内存&lt;/td&gt;
   &lt;td&gt;每 isolate 128 MB&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;构建&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;500 次/月，1 并发，单次超时 20 分钟&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;自定义域名&lt;/td&gt;
   &lt;td&gt;可用&lt;/td&gt;
   &lt;td&gt;每项目 100 个&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;站点文件数&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;20,000 个，单文件 ≤ 25 MiB&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;日志&lt;/td&gt;
   &lt;td&gt;200,000 条/天&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;关键点：Pages 的静态资源请求不占用 Workers 配额；只有 Pages Functions 才计入 Workers 那 10 万次/天。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;数据层：KV + D1 + R2 + Queues + Durable Objects&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;服务&lt;/td&gt;
   &lt;td&gt;免费额度&lt;/td&gt;
   &lt;td&gt;定位&lt;/td&gt;
   &lt;td&gt;核实状态&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Workers KV&lt;/td&gt;
   &lt;td&gt;读 100,000/天、写 1,000/天、删除 1,000/天、列表 1,000/天、存储 1 GB&lt;/td&gt;
   &lt;td&gt;配置、会话、读多写少的缓存&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;D1&lt;/td&gt;
   &lt;td&gt;读 500 万行/天、写 100,000 行/天、总存储 5 GB&lt;/td&gt;
   &lt;td&gt;SQLite 兼容关系库&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;R2&lt;/td&gt;
   &lt;td&gt;存储 10 GB-month/月、A 类操作 100 万/月、B 类操作 1000 万/月、出口流量免费&lt;/td&gt;
   &lt;td&gt;对象存储，对标 S3&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Queues&lt;/td&gt;
   &lt;td&gt;10,000 次操作/天、消息保留 24 小时&lt;/td&gt;
   &lt;td&gt;异步任务、削峰&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Hyperdrive&lt;/td&gt;
   &lt;td&gt;100,000 次数据库查询/天&lt;/td&gt;
   &lt;td&gt;加速访问外部 Postgres/MySQL&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Durable Objects&lt;/td&gt;
   &lt;td&gt;请求 100,000/天、时长 13,000 GB-s/天&lt;/td&gt;
   &lt;td&gt;有状态协作、WebSocket&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Workflows&lt;/td&gt;
   &lt;td&gt;步骤 3,000/天、存储 1 GB-month&lt;/td&gt;
   &lt;td&gt;多步骤任务编排&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;三个容易踩的坑：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;D1 的「读 500 万行」是按扫描行数算，不是返回行数。一条没有索引的 SELECT * FROM t 在 5 万行的表上跑一次就是 5 万行读取，跑 100 次当天额度就没了。索引不是优化项，是配额保护。&lt;/li&gt;
  &lt;li&gt;R2 的 10 GB 免费额度只适用于标准存储。低频访问存储不享受免费额度（官方明确写了 free tier only applies to Standard storage）。&lt;/li&gt;
  &lt;li&gt;Queues 的免费额度实际比看上去小。一条消息完整走完要 3 次操作（写、读、删），1 万次/天大约只够 3000 条消息。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;安全与身份层：Access + Tunnel + Turnstile&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;服务&lt;/td&gt;
   &lt;td&gt;免费额度&lt;/td&gt;
   &lt;td&gt;替代了什么&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Cloudflare Tunnel&lt;/td&gt;
   &lt;td&gt;免费，隧道与路由不限量&lt;/td&gt;
   &lt;td&gt;ngrok、frp、DDNS + 端口转发&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Zero Trust Access&lt;/td&gt;
   &lt;td&gt;50 用户上限，永久免费&lt;/td&gt;
   &lt;td&gt;VPN、跳板机、内网鉴权&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Gateway&lt;/td&gt;
   &lt;td&gt;DNS / HTTP 过滤，免费&lt;/td&gt;
   &lt;td&gt;企业 DNS 过滤&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Turnstile&lt;/td&gt;
   &lt;td&gt;免费，20 个 widget，验证次数不限&lt;/td&gt;
   &lt;td&gt;reCAPTCHA、hCaptcha&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;WARP 客户端&lt;/td&gt;
   &lt;td&gt;个人版免费&lt;/td&gt;
   &lt;td&gt;加密 DNS + 私网接入&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;Tunnel + Access 是免费层里最被低估的组合：它一次性干掉了动态 DNS、端口转发、Nginx 反代配置、certbot 证书续期这四件麻烦事，而且路由器上不需要开任何入站端口。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;观测与效率层&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;服务&lt;/td&gt;
   &lt;td&gt;免费额度&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Web Analytics&lt;/td&gt;
   &lt;td&gt;免费，无 Cookie、无指纹追踪&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Email Routing&lt;/td&gt;
   &lt;td&gt;免费，域名邮箱转发&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Workers AI&lt;/td&gt;
   &lt;td&gt;10,000 Neurons/天&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;AI Gateway&lt;/td&gt;
   &lt;td&gt;核心功能免费&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Registrar&lt;/td&gt;
   &lt;td&gt;按成本价注册，无加价&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Vectorize&lt;/td&gt;
   &lt;td&gt;口径存疑：文档中同时存在「需 Workers Paid」与部分 Free 行，不要作为稳定免费依赖&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Browser Rendering&lt;/td&gt;
   &lt;td&gt;免费额度很小且口径不一，上线前需核对官方文档&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;h2&gt;Cloudflare的使用场景&lt;/h2&gt;
 &lt;h3&gt;场景一：把域名接入 Cloudflare（DNS + TLS + CDN）&lt;/h3&gt;
 &lt;p&gt;这是所有后续操作的前提。根域名要挂 Pages、要开 Tunnel，都必须先让域名成为 Cloudflare 的一个 zone。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;前置条件&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;一个已注册的域名，能修改 nameserver&lt;/li&gt;
  &lt;li&gt;一个 Cloudflare 账号（免费，不需要信用卡）&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;步骤&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;登录cloudflare.com，选择 Add a site，输入根域名。&lt;/li&gt;
  &lt;li&gt;选择 Free 计划。&lt;/li&gt;
  &lt;li&gt;Cloudflare 会自动扫描现有 DNS 记录。这里要做的第一件事是核对扫描结果——自动扫描经常漏掉 MX、TXT、CNAME 记录。改 nameserver 之前，先把原 DNS 服务商的所有记录导出或截图留档。Cloudflare 自动扫描不是 100% 完整，漏一条 MX 记录就意味着域名邮箱立刻失效。&lt;/li&gt;
  &lt;li&gt;去域名注册商处，把 nameserver 改成 Cloudflare 分配的两个地址。&lt;/li&gt;
  &lt;li&gt;等待生效，通常几分钟到 24 小时。&lt;/li&gt;
  &lt;li&gt;进入 SSL/TLS，把加密模式设为 Full (strict)。如果源站没有有效证书，先用 Full，但不要用 Flexible。&lt;/li&gt;
  &lt;li&gt;进入 Speed → Optimization，开启 Brotli 与 Early Hints；Auto Minify 已被弃用，不需要设。&lt;/li&gt;
  &lt;li&gt;验证：&lt;/li&gt;
&lt;/ul&gt;
 &lt;pre&gt;# 确认 nameserver 已切换
dig NS example.com +short

# 确认解析走的是 Cloudflare
dig A example.com +short

# 确认证书与 HTTP 版本
curl -sI https://example.com | head -n 5
&lt;/pre&gt;
 &lt;p&gt;  &lt;strong&gt;三个高频坑&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;现象&lt;/td&gt;
   &lt;td&gt;原因&lt;/td&gt;
   &lt;td&gt;处理&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;522 连接超时&lt;/td&gt;
   &lt;td&gt;源站没有放行 Cloudflare 回源 IP，或回源端口不通&lt;/td&gt;
   &lt;td&gt;放行官方 IP 段，确认 443 可达&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;无限重定向&lt;/td&gt;
   &lt;td&gt;SSL 模式用 Flexible，源站又强制 HTTPS&lt;/td&gt;
   &lt;td&gt;改为 Full (strict)&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;邮件收不到&lt;/td&gt;
   &lt;td&gt;自动扫描漏了 MX 记录&lt;/td&gt;
   &lt;td&gt;手动补 MX，且必须保持 DNS only（灰云）&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;橙色云（代理）与灰色云（仅 DNS）的区别要记住：邮件相关记录必须是灰云，否则收不到信。&lt;/p&gt;
 &lt;h3&gt;场景二：静态站部署到 Pages&lt;/h3&gt;
 &lt;p&gt;Pages 是免费层里最慷慨的产品：无限站点、无限席位、无限静态请求、无限带宽，只有 500 次/月的构建次数是硬限制。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;三种部署方式对比&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;方式&lt;/td&gt;
   &lt;td&gt;适用场景&lt;/td&gt;
   &lt;td&gt;限制&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Git 集成&lt;/td&gt;
   &lt;td&gt;长期维护的项目&lt;/td&gt;
   &lt;td&gt;一次提交触发一次构建，消耗构建配额&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Wrangler CLI&lt;/td&gt;
   &lt;td&gt;本地已有构建流程&lt;/td&gt;
   &lt;td&gt;需要 Node 环境，最灵活&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;拖拽上传&lt;/td&gt;
   &lt;td&gt;一次性静态页&lt;/td&gt;
   &lt;td&gt;最多 1000 个文件，不支持 functions/ 目录编译&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;  &lt;strong&gt;方式一：Git 集成&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;进入 Workers &amp;amp; Pages → Create → Pages → Connect to Git。&lt;/li&gt;
  &lt;li&gt;授权 GitHub / GitLab，选择仓库。&lt;/li&gt;
  &lt;li&gt;配置构建参数：&lt;/li&gt;
&lt;/ul&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;字段&lt;/td&gt;
   &lt;td&gt;说明&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Project name&lt;/td&gt;
   &lt;td&gt;决定 &amp;lt;name&amp;gt;.pages.dev 子域名&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Production branch&lt;/td&gt;
   &lt;td&gt;一般为 main&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Framework preset&lt;/td&gt;
   &lt;td&gt;按框架选择，纯静态 HTML 选 None&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Build command&lt;/td&gt;
   &lt;td&gt;例如 npm run build&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Build output directory&lt;/td&gt;
   &lt;td&gt;例如 dist / out / public&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Root directory&lt;/td&gt;
   &lt;td&gt;monorepo 必填&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Environment variables&lt;/td&gt;
   &lt;td&gt;敏感配置与 NODE_VERSION&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;ul&gt;
  &lt;li&gt;保存并部署。之后每次 push 会自动构建，每个 PR 会生成独立的预览链接。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;方式二：Wrangler CLI&lt;/strong&gt;&lt;/p&gt;
 &lt;pre&gt;npm install -g wrangler
wrangler login
npm run build
wrangler pages project create my-site
wrangler pages deploy dist
&lt;/pre&gt;
 &lt;p&gt;注意 wrangler pages deploy 消耗的是构建配额之外的直接上传通道，但仍受每站点 20,000 文件上限约束。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;三个构建期陷阱&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;报错&lt;/td&gt;
   &lt;td&gt;根因&lt;/td&gt;
   &lt;td&gt;修复&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;npm ci 直接失败&lt;/td&gt;
   &lt;td&gt;lockfile 与 package.json 不同步&lt;/td&gt;
   &lt;td&gt;本地重跑 npm install 提交新 lockfile&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Failed to load SWC binary for linux/x64&lt;/td&gt;
   &lt;td&gt;macOS 生成的 lockfile 只含 darwin 平台可选依赖&lt;/td&gt;
   &lt;td&gt;在 Linux 环境重新生成 lockfile，或显式加 @next/swc-linux-x64-gnu 到 optionalDependencies&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;修完依赖仍报同样错&lt;/td&gt;
   &lt;td&gt;构建缓存复用了旧的 node_modules&lt;/td&gt;
   &lt;td&gt;加 postinstall 兜底，或关闭构建缓存&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;这三条都是同一个根源：构建发生在 Cloudflare 的 Linux 容器里，而你的 lockfile 是本地平台生成的。&lt;/p&gt;
 &lt;p&gt;不要用 &amp;lt;project&amp;gt;.pages.dev 裸域名面向国内用户。该域名段长期存在 DNS 污染与针对性阻断的社区实测反馈，属于不稳定而非完全不可用。绑自定义域名能明显改善，最稳的做法是前置一家国内可用的 CDN。&lt;/p&gt;
 &lt;h3&gt;场景三：Tunnel + Access 暴露内网服务&lt;/h3&gt;
 &lt;p&gt;替代方案是 DDNS + 端口转发 + Nginx 反代 + certbot 四件套。用 Cloudflare 只需要一条出站连接。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;架构&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;本地服务 → cloudflared 出站连接 → Cloudflare 边缘 → 用户 路由器上不需要任何入站端口。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;步骤&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;安装 cloudflared（Windows 用 winget 或直接下载 exe）。&lt;/li&gt;
  &lt;li&gt;登录并授权：cloudflared tunnel login&lt;/li&gt;
  &lt;li&gt;创建隧道：cloudflared tunnel create homelab&lt;/li&gt;
  &lt;li&gt;配置 DNS 路由：cloudflared tunnel route dns homelab nas.example.com&lt;/li&gt;
  &lt;li&gt;写配置文件yml：&lt;/li&gt;
&lt;/ul&gt;
 &lt;pre&gt;tunnel: homelab
credentials-file: C:/Users/you/.cloudflared/&amp;lt;tunnel-id&amp;gt;.json
ingress:
  - hostname: nas.example.com
    service: http://localhost:5000
  - hostname: grafana.example.com
    service: http://localhost:3000
  - service: http_status:404
&lt;/pre&gt;
 &lt;ul&gt;
  &lt;li&gt;前台验证：cloudflared tunnel run homelab&lt;/li&gt;
  &lt;li&gt;确认可访问后，注册为系统服务：cloudflared service install&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;加上 Access 鉴权&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;隧道本身只是把服务接上公网，任何知道域名的人都能访问，所以鉴权这一步不能省。&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;进入 Zero Trust → Access → Applications → Add an application → Self-hosted。&lt;/li&gt;
  &lt;li&gt;填入受保护的域名。&lt;/li&gt;
  &lt;li&gt;创建策略：Action 选 Allow，Include 选 Emails 或 GitHub 组织，填入允许的账号。&lt;/li&gt;
  &lt;li&gt;保存。免费层支持 50 个用户，个人和家庭场景完全够用。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;验证方式：用未授权的浏览器或无痕窗口访问，应该被拦截到登录页。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;注意点&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;yml 里最后一条 service: http_status:404 是兜底规则，不能省，否则未匹配的 hostname 会报配置错误。&lt;/li&gt;
  &lt;li&gt;隧道服务运行在后台时，日志不会自动输出，排查问题先用前台模式跑。&lt;/li&gt;
  &lt;li&gt;Access 的免费层日志保留期只有 24 小时。&lt;/li&gt;
&lt;/ul&gt;
 &lt;h3&gt;场景四：免费域名邮箱（Email Routing）&lt;/h3&gt;
 &lt;p&gt;免费层能把 hello@yourdomain.com 转发到你的常用邮箱，但它不是邮箱托管——只能收，不能发。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;步骤&lt;/strong&gt;&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;进入 Email → Email Routing，点击启用。&lt;/li&gt;
  &lt;li&gt;Cloudflare 会自动添加所需 MX 与 TXT（SPF）记录，确认添加即可。&lt;/li&gt;
  &lt;li&gt;添加目标邮箱，去该邮箱点击验证链接。&lt;/li&gt;
  &lt;li&gt;创建自定义地址：hello@yourdomain.com → 转发到目标邮箱。&lt;/li&gt;
  &lt;li&gt;可选：启用 Catch-all 规则，把所有未匹配地址统一转发。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;strong&gt;限制&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;项目&lt;/td&gt;
   &lt;td&gt;说明&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;发信&lt;/td&gt;
   &lt;td&gt;不支持，需要 SMTP 服务商&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;邮件内容&lt;/td&gt;
   &lt;td&gt;Cloudflare 声明不存储、不访问邮件正文&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;已有 MX 记录&lt;/td&gt;
   &lt;td&gt;会冲突，启用前需确认没有现存的邮箱服务&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;如果需要在域名邮箱里回复，方案是转发到 Gmail 后在 Gmail 里添加「以别名发送」，需要该 SMTP 服务商（如 Zoho Mail 免费版）配合。&lt;/p&gt;
 &lt;h3&gt;场景五：纯免费的全栈小应用&lt;/h3&gt;
 &lt;p&gt;把 Workers + D1 + R2 组合起来，可以搭一个没有固定成本的小后端。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;初始化&lt;/strong&gt;&lt;/p&gt;
 &lt;pre&gt;npm create cloudflare@latest my-api
cd my-api
wrangler d1 create my-db
&lt;/pre&gt;
 &lt;p&gt;  &lt;strong&gt;绑定配置 wrangler.jsonc&lt;/strong&gt;&lt;/p&gt;
 &lt;pre&gt;{
  &amp;quot;name&amp;quot;: &amp;quot;my-api&amp;quot;,
  &amp;quot;main&amp;quot;: &amp;quot;src/index.ts&amp;quot;,
  &amp;quot;compatibility_date&amp;quot;: &amp;quot;2026-09-01&amp;quot;,
  &amp;quot;d1_databases&amp;quot;: [
    { &amp;quot;binding&amp;quot;: &amp;quot;DB&amp;quot;, &amp;quot;database_name&amp;quot;: &amp;quot;my-db&amp;quot;, &amp;quot;database_id&amp;quot;: &amp;quot;&amp;lt;上一步返回的 id&amp;gt;&amp;quot; }
  ],
  &amp;quot;r2_buckets&amp;quot;: [
    { &amp;quot;binding&amp;quot;: &amp;quot;BUCKET&amp;quot;, &amp;quot;bucket_name&amp;quot;: &amp;quot;my-files&amp;quot; }
  ]
}
&lt;/pre&gt;
 &lt;p&gt;  &lt;strong&gt;示例代码 src/index.ts&lt;/strong&gt;&lt;/p&gt;
 &lt;pre&gt;export interface Env {
  DB: D1Database;
  BUCKET: R2Bucket;
}

export default {
  async fetch(request: Request, env: Env): Promise&amp;lt;Response&amp;gt; {
    const url = new URL(request.url);

    if (url.pathname === &amp;quot;/items&amp;quot; &amp;amp;&amp;amp; request.method === &amp;quot;GET&amp;quot;) {
      const { results } = await env.DB
        .prepare(&amp;quot;SELECT id, name FROM items WHERE id = ?1&amp;quot;)
        .bind(Number(url.searchParams.get(&amp;quot;id&amp;quot;)))
        .all();
      return Response.json(results);
    }

    if (url.pathname === &amp;quot;/upload&amp;quot; &amp;amp;&amp;amp; request.method === &amp;quot;PUT&amp;quot;) {
      const key = crypto.randomUUID();
      await env.BUCKET.put(key, request.body);
      return Response.json({ key });
    }

    return new Response(&amp;quot;Not Found&amp;quot;, { status: 404 });
  },
};
&lt;/pre&gt;
 &lt;p&gt;  &lt;strong&gt;建表与部署&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;注意示例里 WHERE id = ?1 命中了主键，只读 1 行。如果换成无索引列的条件查询，读取行数会按扫描量暴涨，直接吃掉 D1 的每日配额。这是免费层里最容易被忽略的隐形消耗。&lt;/p&gt;
 &lt;pre&gt;wrangler d1 execute my-db --remote --command \
  &amp;quot;CREATE TABLE items (id INTEGER PRIMARY KEY, name TEXT NOT NULL);&amp;quot;

wrangler d1 execute my-db --remote --command \
  &amp;quot;CREATE INDEX idx_items_id ON items(id);&amp;quot;

wrangler deploy
&lt;/pre&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;组件&lt;/td&gt;
   &lt;td&gt;免费额度&lt;/td&gt;
   &lt;td&gt;这个应用的实际消耗&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Workers 请求&lt;/td&gt;
   &lt;td&gt;10 万/天&lt;/td&gt;
   &lt;td&gt;按访问量计&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;D1 读取&lt;/td&gt;
   &lt;td&gt;500 万行/天&lt;/td&gt;
   &lt;td&gt;命中索引时每查询 1 行&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;D1 写入&lt;/td&gt;
   &lt;td&gt;10 万行/天&lt;/td&gt;
   &lt;td&gt;每次新增 1 行&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;R2 存储&lt;/td&gt;
   &lt;td&gt;10 GB&lt;/td&gt;
   &lt;td&gt;按文件大小计&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;R2 出口&lt;/td&gt;
   &lt;td&gt;免费&lt;/td&gt;
   &lt;td&gt;下载不产生费用&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;只要按索引查询、上传文件控制在 10 GB 内、日请求低于 10 万次，这个应用的月成本是 0。&lt;/p&gt;
 &lt;h2&gt;免费额度总表&lt;/h2&gt;
 &lt;table width="812"&gt;

  &lt;tr&gt;
   &lt;td&gt;类别&lt;/td&gt;
   &lt;td&gt;服务&lt;/td&gt;
   &lt;td&gt;免费额度&lt;/td&gt;
   &lt;td&gt;核实状态&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;域名&lt;/td&gt;
   &lt;td&gt;DNS / TLS / DDoS&lt;/td&gt;
   &lt;td&gt;无实质限制&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;域名&lt;/td&gt;
   &lt;td&gt;Registrar&lt;/td&gt;
   &lt;td&gt;成本价注册，无加价&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;托管&lt;/td&gt;
   &lt;td&gt;Pages&lt;/td&gt;
   &lt;td&gt;无限带宽、无限请求、500 构建/月、100 域名/项目&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;托管&lt;/td&gt;
   &lt;td&gt;Workers&lt;/td&gt;
   &lt;td&gt;10 万请求/天、10ms CPU、100 脚本&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;KV&lt;/td&gt;
   &lt;td&gt;读 10 万/天、写 1000/天、1 GB&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;D1&lt;/td&gt;
   &lt;td&gt;读 500 万行/天、写 10 万行/天、5 GB&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;R2&lt;/td&gt;
   &lt;td&gt;10 GB-month、A 类 100 万/月、B 类 1000 万/月、出口免费&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;Queues&lt;/td&gt;
   &lt;td&gt;10,000 操作/天、保留 24h&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;Hyperdrive&lt;/td&gt;
   &lt;td&gt;10 万查询/天&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;Durable Objects&lt;/td&gt;
   &lt;td&gt;10 万请求/天、13,000 GB-s/天&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据&lt;/td&gt;
   &lt;td&gt;Workflows&lt;/td&gt;
   &lt;td&gt;3,000 步骤/天、1 GB-month&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;安全&lt;/td&gt;
   &lt;td&gt;WAF&lt;/td&gt;
   &lt;td&gt;自定义规则 + 免费托管规则集&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;安全&lt;/td&gt;
   &lt;td&gt;速率限制&lt;/td&gt;
   &lt;td&gt;1 条规则&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;安全&lt;/td&gt;
   &lt;td&gt;Turnstile&lt;/td&gt;
   &lt;td&gt;20 widget，验证次数不限&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;安全&lt;/td&gt;
   &lt;td&gt;Zero Trust Access&lt;/td&gt;
   &lt;td&gt;50 用户&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;安全&lt;/td&gt;
   &lt;td&gt;Tunnel&lt;/td&gt;
   &lt;td&gt;隧道与路由不限量&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;观测&lt;/td&gt;
   &lt;td&gt;Web Analytics&lt;/td&gt;
   &lt;td&gt;免费&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;邮件&lt;/td&gt;
   &lt;td&gt;Email Routing&lt;/td&gt;
   &lt;td&gt;免费转发&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;AI&lt;/td&gt;
   &lt;td&gt;Workers AI&lt;/td&gt;
   &lt;td&gt;10,000 Neurons/天&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;AI&lt;/td&gt;
   &lt;td&gt;Vectorize&lt;/td&gt;
   &lt;td&gt;口径存在冲突，勿依赖&lt;/td&gt;
   &lt;td&gt;口径存疑&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;AI&lt;/td&gt;
   &lt;td&gt;Browser Rendering&lt;/td&gt;
   &lt;td&gt;免费额度小且口径不一&lt;/td&gt;
   &lt;td&gt;未核实&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;媒体&lt;/td&gt;
   &lt;td&gt;Images / Stream&lt;/td&gt;
   &lt;td&gt;不属于免费核心能力&lt;/td&gt;
   &lt;td&gt;已核实&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;h2&gt;风险与坑&lt;/h2&gt;
 &lt;p&gt;  &lt;strong&gt;免费层没有 SLA&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;免费计划不提供服务等级承诺，也没有工单支持通道。把关键业务押在免费层上，等于接受「可能突然不可用且无人负责」。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;免费额度是硬中断&lt;/strong&gt;&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;错误码&lt;/td&gt;
   &lt;td&gt;含义&lt;/td&gt;
   &lt;td&gt;表现&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1015&lt;/td&gt;
   &lt;td&gt;触发突发速率限制&lt;/td&gt;
   &lt;td&gt;1000 请求/分钟以上，被拦截&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1027&lt;/td&gt;
   &lt;td&gt;Worker 因超日限额被暂停&lt;/td&gt;
   &lt;td&gt;路由默认 fail open（绕过 Worker），可改为 fail closed&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;1016&lt;/td&gt;
   &lt;td&gt;源站 DNS 解析错误&lt;/td&gt;
   &lt;td&gt;回源域名解析失败&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;522&lt;/td&gt;
   &lt;td&gt;回源连接超时&lt;/td&gt;
   &lt;td&gt;源站未放行或端口不通&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;fail open 与 fail closed 的选择很关键：安全类 Worker（鉴权、校验）应该设成 fail closed，否则额度耗尽时请求会直接穿透到源站，等于安全策略失效。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;CDN 代理非 HTML 内容有条款限制&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;Cloudflare 服务条款第 2.8 节限制用 CDN 代理大比例的流媒体与部分非 HTML 内容。个人博客、文档站、图片站没有风险，但把它当免费视频 CDN 用会违反条款。视频走 R2 或 Stream 是合规路径。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;构建环境与本地环境不一致&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;前面提到的 lockfile 平台差异问题，本质是所有云端构建的通病。对策是锁定 NODE_VERSION、显式声明平台相关可选依赖、必要时关闭构建缓存。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;新账号的隐性限制&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;新注册账号在最初一段时间创建项目可能受限（防滥用机制）。如果第一天就报错，通常不是配置问题，等一等即可。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;国内访问不稳定&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;pages.dev 与 workers.dev 裸域名长期存在解析异常反馈。生产域名必须绑自定义域名，主要面向国内用户时建议前置国内可用的 CDN。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;容易误判的免费项&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;以下几项常被当成免费能力，实际不是或口径不明：Images、Stream、Argo Smart Routing、Load Balancing、高级 Bot Management、Containers、Vectorize。规划架构时不要把它们算进免费方案。&lt;/p&gt;
 &lt;h2&gt;什么时候该付那 5 美元&lt;/h2&gt;
 &lt;p&gt;Workers Paid 是每月 5 美元，包含 1000 万请求、3000 万 CPU 毫秒、1000 万 KV 读取。判断标准很清晰：&lt;/p&gt;
 &lt;table&gt;

  &lt;tr&gt;
   &lt;td&gt;信号&lt;/td&gt;
   &lt;td&gt;说明&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;日请求逼近 10 万&lt;/td&gt;
   &lt;td&gt;免费层没有弹性，超了直接 1027&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;单次任务 CPU 超 10ms&lt;/td&gt;
   &lt;td&gt;免费层的 CPU 上限锁死了计算类任务&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;需要 Durable Objects 完整能力&lt;/td&gt;
   &lt;td&gt;WebSocket 与有状态场景&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;需要有意义的技术支持&lt;/td&gt;
   &lt;td&gt;免费层只有社区论坛&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;业务不能接受无 SLA&lt;/td&gt;
   &lt;td&gt;这是最实际的付费理由&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;反过来说，如果只是博客、文档站、家庭内网、小工具 API，免费层可以长期稳定运行，没有必须付费的理由。&lt;/p&gt;
 &lt;h2&gt;参考资料&lt;/h2&gt;
 &lt;ul&gt;
  &lt;li&gt;Cloudflare 开发者平台定价：https://www.cloudflare.com/plans/developer-platform&lt;/li&gt;
  &lt;li&gt;Workers 定价与限制：https://developers.cloudflare.com/workers/platform/pricing&lt;/li&gt;
  &lt;li&gt;Workers 限制文档：https://developers.cloudflare.com/workers/platform/limits&lt;/li&gt;
  &lt;li&gt;Pages 定价：https://pages.cloudflare.com&lt;/li&gt;
  &lt;li&gt;R2 定价：https://developers.cloudflare.com/r2/pricing&lt;/li&gt;
  &lt;li&gt;D1 产品页：https://www.cloudflare.com/developer-platform/products/d1&lt;/li&gt;
  &lt;li&gt;Zero Trust 定价：   &lt;a href="https://www.cloudflare.com/plans/zero-trust-services"&gt;https://www.cloudflare.com/plans/zero-trust-services&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
 &lt;div&gt;

  &lt;strong&gt;相关文章:&lt;/strong&gt;  &lt;ol&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/cdn/" rel="bookmark" title="&amp;#20869;&amp;#23481;&amp;#20998;&amp;#21457;&amp;#32593;&amp;#32476;CDN"&gt;内容分发网络CDN&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/dns-server/" rel="bookmark" title="&amp;#22914;&amp;#20309;&amp;#25645;&amp;#24314;DNS&amp;#26381;&amp;#21153;"&gt;如何搭建DNS服务&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/causalml/" rel="bookmark" title="&amp;#24320;&amp;#28304;&amp;#22240;&amp;#26524;&amp;#25512;&amp;#26029;&amp;#24211;CausalML"&gt;开源因果推断库CausalML&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>器→工具 工具软件 术→技巧 运维 免费</category>
      <guid isPermaLink="true">https://itindex.net/detail/63282-cloudflare-%E5%85%8D%E8%B4%B9-%E6%9C%8D%E5%8A%A1</guid>
      <pubDate>Mon, 14 Sep 2026 21:22:19 CST</pubDate>
    </item>
    <item>
      <title>基于 AI 生成 HTML 动态页面，然后录视频，是否是制作短视频的一个可行方案</title>
      <link>https://itindex.net/detail/63281-ai-html-%E5%8A%A8%E6%80%81%E9%A1%B5%E9%9D%A2</link>
      <description>&lt;h2&gt;AI Agent + HTML转视频框架（全自动）&lt;/h2&gt;

 &lt;p&gt;这是当前更前沿的“Vibe Motion”趋势，将录屏步骤也自动化了。代表性开源框架包括：&lt;/p&gt;

 &lt;ul&gt;
  &lt;li&gt;HyperFrames：由 HeyGen 开源，专为 AI Agent 设计。你只需用自然语言描述需求，Agent 就会编写 HTML+CSS+GSAP 动画代码，然后框架通过无头浏览器逐帧捕获，用 FFmpeg 直接渲染成 MP4。   &lt;br /&gt;
&lt;/li&gt;
  &lt;li&gt;html-video：Open Design 团队出品，基于 HyperFrames 构建。它提供21套预设模板（如数据可视化、电影感标题），支持粘贴文章链接自动生成视频，进一步降低了使用门槛。   &lt;br /&gt;
&lt;/li&gt;
&lt;/ul&gt;

 &lt;h2&gt;HyperFrames&lt;/h2&gt;

 &lt;p&gt;https://github.com/heygen-com/hyperframes&lt;/p&gt;

 &lt;p&gt;HyperFrames 是一个开源框架，可将 HTML、CSS、媒体内容以及可播放的动画转换为格式固定的 MP4 视频。它可以通过命令行工具在本地使用，也可以与具备相应功能的 AI 编程工具配合使用。此外，它还可以作为托管式内容创作流程中的渲染核心组件来使用。&lt;/p&gt;

 &lt;h2&gt;html-video&lt;/h2&gt;

 &lt;p&gt;https://github.com/nexu-io/html-video&lt;/p&gt;

 &lt;p&gt;HTML 内容可以直接转换成视频——而且是在你的笔记本电脑上完成。你可以使用各种编程工具来辅助这个过程：Open Design、Windsurf CLI、Trae CLI、Claude Code、Cursor、Codex、Gemini、Grok、Qwen、OpenCode、Copilot、Aider、Hermes，或者 Anthropic API。你只需描述一下想要制作的视频内容，或者粘贴相关文章的链接或 GitHub 代码库的地址，这些工具就会把内容转换成多帧动画视频。最后，这些视频会被直接保存为 MP4 格式，保存在你的电脑上。整个过程非常简单：只需使用相应的工具，无需额外付费，也不会受到任何软件或服务的限制。Apache-2.0 许可证，完全免费使用，无任何附加条件。&lt;/p&gt;

 &lt;h2&gt;选哪个方案&lt;/h2&gt;

 &lt;p&gt;我看了一下这俩个 github 项目，HyperFrames 无论是 star 数量，还是更新频率，及 commit 更新数量，都远远优于 html-video。所以，我还是优先尝试 HyperFrames 吧。&lt;/p&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63281-ai-html-%E5%8A%A8%E6%80%81%E9%A1%B5%E9%9D%A2</guid>
      <pubDate>Sun, 13 Sep 2026 10:30:03 CST</pubDate>
    </item>
    <item>
      <title>编程 Agent 连接本机大模型的注意点</title>
      <link>https://itindex.net/detail/63280-%E7%BC%96%E7%A8%8B-agent-%E6%A8%A1%E5%9E%8B</link>
      <description>&lt;h1&gt;如果你曾经考虑过把编程 Harness 发出的 API 调用替换成本地模型，那么结果很可能令人失望。你可能快速跑了一下   &lt;code&gt;llama-bench&lt;/code&gt;，以为自己会得到一个“每秒 X token”的结果。但这个基准测试无法告诉你  &lt;strong&gt;实际使用这些 Harness 时的体验&lt;/strong&gt;。&lt;/h1&gt; &lt;p&gt;不同 Harness 的开发体验差异非常大。你可能会坐在那里等几分钟才看到模型开始响应；或者刚开始运行得不错，中途却突然卡住。&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;“是谁把 Prefill 的时间全吃光了？”&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;这并不是你的错：  &lt;strong&gt;大多数编程 Harness 从设计之初就没有考虑本地模型。&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;这并不是要批评其他团队的优秀工作。我只是想可靠地测量：  &lt;strong&gt;当你把数据中心换成本地主机（localhost）之后，实际会发生什么。&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;需要说明的是，我一直在折腾   &lt;strong&gt;chad&lt;/strong&gt;，这是一个专门针对 Apple Silicon 上 Qwen 3.8 27B 优化的 coding harness。&lt;/p&gt; &lt;h2&gt;Laptop Physics：笔记本的物理限制&lt;/h2&gt; &lt;p&gt;大多数 Harness 会从三个方面与本地模型“作对”。&lt;/p&gt; &lt;h3&gt;1. 巨大的 System Prompt 和 Tool Schema&lt;/h3&gt; &lt;p&gt;假设你的笔记本读取速度是   &lt;strong&gt;90 tokens/s&lt;/strong&gt;，生成速度大约   &lt;strong&gt;10 tokens/s&lt;/strong&gt;。&lt;/p&gt; &lt;p&gt;这分别对应编程 Harness 每一轮中的   &lt;strong&gt;Prefill 和 Generation&lt;/strong&gt;。&lt;/p&gt; &lt;p&gt;按照这个速度：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;每增加 1,000 个 Prompt Token，就意味着模型开始输出前，你要盯着光标等待约 11 秒。&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;而在模型真正开始干活之前，它必须先读取：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;整个 System Prompt&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;所有加载的 Tool Schema&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;例如：&lt;/p&gt; &lt;table&gt;  &lt;tr&gt;   &lt;th&gt;Harness&lt;/th&gt;   &lt;th align="right"&gt;Prompt Token&lt;/th&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;pi&lt;/td&gt;   &lt;td align="right"&gt;2,008&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;OpenCode&lt;/td&gt;   &lt;td align="right"&gt;18,046&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt; &lt;p&gt;在数据中心 GPU 上，Prefill 可能达到   &lt;strong&gt;10k+ tokens/s&lt;/strong&gt;，因此：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;pi：约     &lt;strong&gt;0.2 秒&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;OpenCode：约     &lt;strong&gt;1.8 秒&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;几乎感觉不到区别。&lt;/p&gt; &lt;p&gt;但在你的笔记本上：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;pi：约     &lt;strong&gt;22 秒&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;OpenCode：约     &lt;strong&gt;226 秒&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;也就是  &lt;strong&gt;将近 4 分钟&lt;/strong&gt;。&lt;/p&gt; &lt;p&gt;这已经完全无法接受了。&lt;/p&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h3&gt;2. 更小的有效 Context Window&lt;/h3&gt; &lt;p&gt;LLM 处理完 System Prompt 后，真正能够用于工作的 Context Window 已经减少了。&lt;/p&gt; &lt;p&gt;你的上下文窗口其实比你想象得小。&lt;/p&gt; &lt;p&gt;具体剩余多少，取决于：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;你的内存大小&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;模型权重占用&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV Cache&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;其他运行时开销&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;对于一台配置合理的笔记本，可以粗略认为还有   &lt;strong&gt;32,000 tokens&lt;/strong&gt; 可用。&lt;/p&gt; &lt;p&gt;那么：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;pi：&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;System Prompt 只占 2,008 tokens。&lt;/p&gt; &lt;p&gt;大约还有   &lt;strong&gt;94%&lt;/strong&gt; 的 Context 可以真正用于工作。&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;OpenCode：&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;一开始就消耗了 18,046 tokens。&lt;/p&gt; &lt;p&gt;真正用于工作的 Context 只剩下约   &lt;strong&gt;44%&lt;/strong&gt;。&lt;/p&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h3&gt;3. 大量 Side Request&lt;/h3&gt; &lt;p&gt;传统的：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;本地客户端 → 远程服务器&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;架构下，Harness 可以随意向数据中心发起各种额外请求。&lt;/p&gt; &lt;p&gt;但本地模型不同：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;你的笔记本既是 Client，又是 Server。&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;最好的情况是这些 Side Request 排队等待。&lt;/p&gt; &lt;p&gt;最糟糕的情况是：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;它们不断触发重复 Prefill。&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;在测试的 24 个任务中：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;OpenCode：33 次 Side Request&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;Crush：51 次&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;dsh：24 次&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;而且几乎都与 Agent Turn 重叠。&lt;/p&gt; &lt;p&gt;结果就是：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;OpenCode：GPU 忙碌时间达到实际时间的     &lt;strong&gt;125%&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;Crush：    &lt;strong&gt;114%&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;也就是说：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;一块 GPU 上同时跑了两个请求。&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h1&gt;Harness 测试结果&lt;/h1&gt; &lt;p&gt;作者使用   &lt;strong&gt;9 个 Coding Harness&lt;/strong&gt;，测试了   &lt;strong&gt;8 个 Exercism 编程任务&lt;/strong&gt;。&lt;/p&gt; &lt;p&gt;每个任务都使用：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;相同的一句话 Prompt&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;Auto-approve 模式&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;同一台     &lt;strong&gt;M4 MacBook Pro&lt;/strong&gt;&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;24GB 内存&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;macOS 26.6.2&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;Qwen 3.8 27B&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;3-bit 量化&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;llama.cpp build 10470&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;所有 Harness 共用同一个 llama-server。&lt;/p&gt; &lt;p&gt;采样参数统一为：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;temperature = 1.0&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;top_k = 20&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;top_p = 0.95&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;min_p = 0.05&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;Context Cache 统一：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;32,768 tokens / 4 slots&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;其中：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;    &lt;strong&gt;wait before 1st token&lt;/strong&gt;：第一次生成 Token 前的等待时间，也就是 System Prompt + Tool Schema + First Request 的 Prefill 时间&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;    &lt;strong&gt;wait / later turn&lt;/strong&gt;：之后每一轮的等待时间&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;    &lt;strong&gt;cache reuse&lt;/strong&gt;：后续 Turn 有多少比例来自 Prefix Cache&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;    &lt;strong&gt;experienced tokens/s&lt;/strong&gt;：真正用户体验到的速度，包括 Prefill、Tool 等全部时间&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;    &lt;strong&gt;pass&lt;/strong&gt;：Exercism 任务是否通过，作者特别强调不要把它当成模型能力排名&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;table&gt;  &lt;tr&gt;   &lt;th&gt;Harness&lt;/th&gt;   &lt;th align="right"&gt;Tools&lt;/th&gt;   &lt;th align="right"&gt;首轮 Prompt&lt;/th&gt;   &lt;th align="right"&gt;首 Token 等待&lt;/th&gt;   &lt;th align="right"&gt;后续 Turn&lt;/th&gt;   &lt;th align="right"&gt;Cache&lt;/th&gt;   &lt;th align="right"&gt;实际 tok/s&lt;/th&gt;   &lt;th align="right"&gt;通过&lt;/th&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;mini-swe-agent&lt;/td&gt;   &lt;td align="right"&gt;1&lt;/td&gt;   &lt;td align="right"&gt;1,171&lt;/td&gt;   &lt;td align="right"&gt;12.2s&lt;/td&gt;   &lt;td align="right"&gt;3.6s&lt;/td&gt;   &lt;td align="right"&gt;96%&lt;/td&gt;   &lt;td align="right"&gt;8.0&lt;/td&gt;   &lt;td align="right"&gt;11/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;pi&lt;/td&gt;   &lt;td align="right"&gt;4&lt;/td&gt;   &lt;td align="right"&gt;2,008&lt;/td&gt;   &lt;td align="right"&gt;21.6s&lt;/td&gt;   &lt;td align="right"&gt;1.3s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;8.1&lt;/td&gt;   &lt;td align="right"&gt;19/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;cline&lt;/td&gt;   &lt;td align="right"&gt;26&lt;/td&gt;   &lt;td align="right"&gt;5,876&lt;/td&gt;   &lt;td align="right"&gt;64.1s&lt;/td&gt;   &lt;td align="right"&gt;9.9s&lt;/td&gt;   &lt;td align="right"&gt;94%&lt;/td&gt;   &lt;td align="right"&gt;7.3&lt;/td&gt;   &lt;td align="right"&gt;17/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;codex&lt;/td&gt;   &lt;td align="right"&gt;10&lt;/td&gt;   &lt;td align="right"&gt;7,804&lt;/td&gt;   &lt;td align="right"&gt;87.8s&lt;/td&gt;   &lt;td align="right"&gt;9.6s&lt;/td&gt;   &lt;td align="right"&gt;94%&lt;/td&gt;   &lt;td align="right"&gt;6.9&lt;/td&gt;   &lt;td align="right"&gt;19/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;dsh&lt;/td&gt;   &lt;td align="right"&gt;25&lt;/td&gt;   &lt;td align="right"&gt;8,052&lt;/td&gt;   &lt;td align="right"&gt;94.4s&lt;/td&gt;   &lt;td align="right"&gt;2.2s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;7.2&lt;/td&gt;   &lt;td align="right"&gt;18/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;goose&lt;/td&gt;   &lt;td align="right"&gt;18&lt;/td&gt;   &lt;td align="right"&gt;9,617&lt;/td&gt;   &lt;td align="right"&gt;110.3s&lt;/td&gt;   &lt;td align="right"&gt;1.0s&lt;/td&gt;   &lt;td align="right"&gt;100%&lt;/td&gt;   &lt;td align="right"&gt;8.0&lt;/td&gt;   &lt;td align="right"&gt;22/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;crush&lt;/td&gt;   &lt;td align="right"&gt;26&lt;/td&gt;   &lt;td align="right"&gt;16,263&lt;/td&gt;   &lt;td align="right"&gt;199.8s&lt;/td&gt;   &lt;td align="right"&gt;1.8s&lt;/td&gt;   &lt;td align="right"&gt;100%&lt;/td&gt;   &lt;td align="right"&gt;5.8&lt;/td&gt;   &lt;td align="right"&gt;18/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;opencode&lt;/td&gt;   &lt;td align="right"&gt;10&lt;/td&gt;   &lt;td align="right"&gt;18,046&lt;/td&gt;   &lt;td align="right"&gt;225.7s&lt;/td&gt;   &lt;td align="right"&gt;4.6s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;5.7&lt;/td&gt;   &lt;td align="right"&gt;15/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;chad（llama.cpp）&lt;/td&gt;   &lt;td align="right"&gt;5&lt;/td&gt;   &lt;td align="right"&gt;2,563&lt;/td&gt;   &lt;td align="right"&gt;25.6s&lt;/td&gt;   &lt;td align="right"&gt;0.8s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;7.9&lt;/td&gt;   &lt;td align="right"&gt;24/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;chad（MLX serial）&lt;/td&gt;   &lt;td align="right"&gt;5&lt;/td&gt;   &lt;td align="right"&gt;2,566&lt;/td&gt;   &lt;td align="right"&gt;4.7s&lt;/td&gt;   &lt;td align="right"&gt;1.0s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;12.4&lt;/td&gt;   &lt;td align="right"&gt;21/24&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;chad（MLX dflash2）&lt;/td&gt;   &lt;td align="right"&gt;5&lt;/td&gt;   &lt;td align="right"&gt;2,562&lt;/td&gt;   &lt;td align="right"&gt;4.6s&lt;/td&gt;   &lt;td align="right"&gt;0.9s&lt;/td&gt;   &lt;td align="right"&gt;99%&lt;/td&gt;   &lt;td align="right"&gt;17.4&lt;/td&gt;   &lt;td align="right"&gt;22/24&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h1&gt;作者把这些 Harness 分成三类&lt;/h1&gt; &lt;h3&gt;① Lean and Stable：轻量、稳定&lt;/h3&gt; &lt;p&gt;  &lt;strong&gt;pi、mini-swe-agent、chad&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;特点：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;精简 System Prompt&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;96–99% Cache Reuse&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;本地模型和云端模型都能比较好地工作&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;其中 mini-swe-agent 的问题是：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;24 个任务只通过 11 个，并且 Timeout 最多。&lt;/p&gt;&lt;/blockquote&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h3&gt;② Heavy but Disciplined：比较重，但设计合理&lt;/h3&gt; &lt;p&gt;  &lt;strong&gt;dsh、cline、codex、goose&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;它们的问题主要来自：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;较长的 System Prompt&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;大量 Tool Schema&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;但有一个很重要的优点：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;Prefix 是稳定的。&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;所以一旦 Prefill 完成，后面还是比较能够接受的。&lt;/p&gt; &lt;p&gt;作者特别提到 Goose：&lt;/p&gt; &lt;p&gt;1.50.0 版本已经解决了一个问题。&lt;/p&gt; &lt;p&gt;早期版本每一轮都会把一个精确到分钟的 Timestamp 重新写进第一条 User Message，导致 Cache Reuse 从接近 100% 降到了   &lt;strong&gt;78%&lt;/strong&gt;。&lt;/p&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h3&gt;③ Heavy to Start：启动非常重&lt;/h3&gt; &lt;p&gt;  &lt;strong&gt;Crush、OpenCode&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;你会先等：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;3～4 分钟&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;才能看到模型开始工作。&lt;/p&gt; &lt;hr&gt;&lt;/hr&gt; &lt;h1&gt;为什么作者认为 chad 特别适合本地模型？&lt;/h1&gt; &lt;p&gt;作者认为这些 Harness   &lt;strong&gt;并不是设计得不好&lt;/strong&gt;。&lt;/p&gt; &lt;p&gt;例如：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;OpenCode 使用 18K Token 的 System Prompt&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;这对于：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;数据中心 + Frontier Model + 10k+ tokens/s Prefill&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;完全合理。&lt;/p&gt; &lt;p&gt;同样：&lt;/p&gt; &lt;p&gt;  &lt;strong&gt;Crush 使用 26 个 Tool Schema&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;当你拥有：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;200K Context&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;几乎免费的高速 Prefill&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;这也完全没问题。&lt;/p&gt; &lt;p&gt;问题在于：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;这些设计都是建立在“Prefill 几乎免费”的数据中心环境上的。&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;一旦换成：&lt;/p&gt; &lt;blockquote&gt;  &lt;p&gt;   &lt;strong&gt;本地模型 + 笔记本 + 低 Prefill 速度&lt;/strong&gt;&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;原来的设计就会彻底暴露问题。&lt;/p&gt; &lt;p&gt;原文：https://nasutton.notion.site/Nine-coding-harnesses-vs-your-laptop-3d139990182b80d59fa3cf500f0450ba&lt;/p&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63280-%E7%BC%96%E7%A8%8B-agent-%E6%A8%A1%E5%9E%8B</guid>
      <pubDate>Fri, 11 Sep 2026 15:01:51 CST</pubDate>
    </item>
    <item>
      <title>科学家建议冲马桶合盖以减少气凝胶</title>
      <link>https://itindex.net/detail/63279-%E7%A7%91%E5%AD%A6%E5%AE%B6-%E9%A9%AC%E6%A1%B6-%E6%B0%94%E5%87%9D%E8%83%B6</link>
      <description>Flinders 大学的研究人员发现，冲马桶会向周围空气释放气溶胶和生物气溶胶，气溶胶颗粒甚至会进入到成年人的呼吸区，而冲水后气溶胶会在空气中悬浮至少 20 秒。这些发现是基于对 22 项马桶气溶胶研究的分析。结果表明，保持良好的厕所卫生，包括定期清洁马桶及其周围表面，以及使用后洗手，有助于最大限度减少微生物污染和潜在的微生物疾病风险。使用马桶的低冲水模式也有助于最大限度减少气溶胶的产生。充足的通风有助于扩散和清除悬浮的空气颗粒，关闭马桶盖会改变气溶胶的扩散方向，气溶胶会从马桶盖和马桶座之间的缝隙逸出，而不是向上扩散。研究人员建议保持卫生间通风良好，在冲水前盖上马桶盖。
 &lt;p&gt;&lt;/p&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63279-%E7%A7%91%E5%AD%A6%E5%AE%B6-%E9%A9%AC%E6%A1%B6-%E6%B0%94%E5%87%9D%E8%83%B6</guid>
      <pubDate>Tue, 08 Sep 2026 23:16:05 CST</pubDate>
    </item>
    <item>
      <title>人工智能的效率提升可能会让我们失去下一代专家（如何保护人类技能）</title>
      <link>https://itindex.net/detail/63278-%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD-%E6%8F%90%E5%8D%87-%E5%A4%B1%E5%8E%BB</link>
      <description>&lt;p&gt;十多年前，我曾主导设计美国一座核电站的首套全数字控制系统。从设计图纸上看，这套系统堪称完美——它被设计成像现代客机一样自主运行，操作员只需偶尔监控一下，几乎不需要他们干预。然而，我们却做出了一个在注重效率的观察者看来有些倒退的决定：我们故意在系统能够自主执行的程序中保留了一些手动步骤。&lt;/p&gt; &lt;div&gt;&lt;/div&gt; &lt;p&gt;我们当时在解决一个具体问题。一个只负责监督  &lt;a href="https://spectrum.ieee.org/tag/automation" rel="noopener noreferrer" target="_blank"&gt;自动化系统的&lt;/a&gt;操作员，会逐渐失去操作能力。他的手会变得冰冷，对工厂实际运行情况的认知也会变得模糊。然后，自动化系统就会把控制权交还给他。这总是最糟糕的一天，因为自动化系统只有在出现故障或出错时才会停止工作。但到那时，坐在操作台上的人已经好几年没真正操作过这套系统了。手动操作步骤的存在是为了保持操作员的技能水平。这种设计本身就效率低下，而且是故意为之。&lt;/p&gt; &lt;p&gt;那个核电站最终并没有建成。由于美国  &lt;a href="https://spectrum.ieee.org/tag/nuclear-power" target="_blank"&gt;核电行业&lt;/a&gt;复杂的政治和经济因素，这个项目被搁置了，而这些因素与工程设计本身无关。但是，设计理念却超越了项目本身，而且我逐渐意识到，对于如今席卷所有董事会的争论——“当人工智能取代人类的专业知识来构建它们时，人类的专业知识将何去何从？”——这或许是我能提出的最有价值的观点。&lt;/p&gt; &lt;h2&gt;人工智能正在颠覆工程职业发展阶梯&lt;/h2&gt; &lt;p&gt;这些数据已不容忽视。哈佛大学一份  &lt;a href="https://spectrum.ieee.org/tag/harvard" rel="noopener noreferrer" target="_blank"&gt;涵盖&lt;/a&gt;超过28万家美国公司约6500万名员工的工作报告发现，在公司采用  &lt;a href="https://spectrum.ieee.org/tag/generative-ai" rel="noopener noreferrer" target="_blank"&gt;生成式人工智能&lt;/a&gt;后，  &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5425555" target="_blank"&gt;初级员工的就业率在六个季度内比未采用人工智能的公司下降了约9%&lt;/a&gt;，而高级员工的就业率则持续增长。  &lt;a href="https://spectrum.ieee.org/tag/stanford" rel="noopener noreferrer" target="_blank"&gt;斯坦福大学&lt;/a&gt;对ADP工资记录的分析也指向了同样的结果：在人工智能应用最广泛的职业领域，最年轻的员工  &lt;a href="https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/" target="_blank"&gt;在2022年底之后就业率下降，&lt;/a&gt;而他们经验更丰富的同事则保持了就业水平。斯坦福大学的研究人员发现，就业率下降主要集中在人工智能实现工作自动化的领域；而在人工智能仅仅是辅助工作的领域，初级员工的就业率则保持稳定或有所上升。&lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;div&gt;&lt;/div&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;因果关系仍存在争议，坦诚相告至关重要。纽约联邦储备银行的研究人员认为，应届毕业生失业率上升的主要原因  &lt;a href="https://libertystreeteconomics.newyorkfed.org/2026/06/remote-work-leaves-younger-workers-sidelined/" target="_blank"&gt;并非人工智能，而是远程办公&lt;/a&gt;。他们认为，企业不愿雇用缺乏经验、无法进行远程培训和指导的人员。但请注意这些解释的共同之处。无论是人工智能吸收了早期培训，还是远程办公切断了指导关系，两者都描述了同一个失效的机制：学徒制渠道——即专业知识从资深员工传授给初级员工的途径——已经失效。无论如何，“入门级”悄然变成了“需要三年工作经验”。&lt;/p&gt; &lt;p&gt;抛开一切杂音，你会发现一个看似简单却又至关重要的问题：你不可能不先当过初级工程师就成为高级工程师  &lt;a href="https://spectrum.ieee.org/ai-effect-entry-level-jobs" target="_blank"&gt;。&lt;/a&gt;  &lt;em&gt;专业&lt;/em&gt;知识并非唾手可得，而是需要通过失败的构建、毫无进展的调试以及那些“这到底是怎么回事”的困惑时刻来积累的。而如今，人工智能会很乐意让新手免于经历这些。如果让新手免于经历足够多的此类困惑，你就会培养出一批能够纸上操作模型，却始终无法培养出那种能够准确判断模型何时出现灾难性错误的直觉的人才。&lt;/p&gt; &lt;p&gt;大多数评论止步于诊断，或者试图提出将初级岗位流失视为经济问题的政策解决方案。然而，这同时也是一个工程问题，而安全攸关领域已经花费数十年时间研究如何解决这个问题。&lt;/p&gt; &lt;h2&gt;  &lt;a href="https://spectrum.ieee.org/tag/automation-paradox" rel="noopener noreferrer" target="_blank"&gt;航空业关于自动化悖论&lt;/a&gt;的教训  &lt;a href="https://spectrum.ieee.org/tag/automation-paradox" rel="noopener noreferrer" target="_blank"&gt;&lt;/a&gt;&lt;/h2&gt; &lt;p&gt;我的职业生涯始于自动化领域的尖端。毕业后的第一份工作是验证和确认数字喷气发动机控制器中的软件，该控制器能够比任何飞行员更快地决定战斗机发动机的响应。即使在20世纪80年代末，这种核心矛盾也显而易见：在常规情况下，机器的性能优于人类；但在机器无法预料的情况下，人类却是飞机免于灾难的唯一保障。这种矛盾被称为“  &lt;a href="https://spectrum.ieee.org/tag/automation-paradox" target="_blank"&gt;自动化悖论”&lt;/a&gt;，即随着自动化能力的不断提升，操作人员的实践机会越来越少，而留给他们的却只有最棘手的难题。&lt;/p&gt; &lt;p&gt;航空业曾多次付出惨痛的代价，才明白当人为技能在自动化系统和人工操作之间的鸿沟中退化时，会发生什么。最典型的例子就是2009年坠入大西洋的  &lt;a href="https://en.wikipedia.org/wiki/Air_France_Flight_447" target="_blank"&gt;法航447航班。事故的直接原因很普通：结冰的空速传感器向&lt;/a&gt;  &lt;a href="https://spectrum.ieee.org/tag/autopilot" rel="noopener noreferrer" target="_blank"&gt;自动驾驶仪&lt;/a&gt;提供了错误数据，自动驾驶仪按照其设计程序断开连接，将飞机控制权交还给机组人员。接下来发生的并非硬件故障，而是飞行员能力的缺失。原本可以挽回的局面变成了无法挽回的悲剧，因为飞行员们在数千小时的自动驾驶操作下，无法识别高空气动失速并手动驾驶飞机脱困。飞机本身没有问题，但被自动化系统悄然削弱的飞行员训练却失效了。&lt;/p&gt; &lt;p&gt;&lt;/p&gt; &lt;div&gt;&lt;/div&gt; &lt;p&gt;&lt;/p&gt; &lt;p&gt;业界的应对措施颇具启发性，与我们在核电站控制室采取的措施如出一辙。他们并没有拆除自动驾驶系统，而是重新引入了刻意的手动飞行练习。2017年，美国联邦航空  &lt;a href="https://spectrum.ieee.org/tag/faa" rel="noopener noreferrer" target="_blank"&gt;管理局（FAA）&lt;/a&gt;发布了第17007号运营商安全警示——《  &lt;a href="https://www.faa.gov/sites/faa.gov/files/2022-11/SAFO17007.pdf" target="_blank"&gt;手动飞行操作熟练度&lt;/a&gt;》，其中明确指出“手动飞行是其他飞行技术的基础”。该警示正式承认技能衰退本身就是一种危险。一些航空公司修改了操作规程，鼓励在天气状况良好的情况下进行初始爬升和初始下降阶段的手动飞行，他们明知会牺牲少量  &lt;a href="https://spectrum.ieee.org/tag/fuel-efficiency" rel="noopener noreferrer" target="_blank"&gt;燃油效率&lt;/a&gt;，却依然努力保持机组人员的飞行技能。这种权衡取舍至关重要。一个看似完美却培养出不称职操作员的系统，根本算不上完美。它只是把故障模式转移到了电子表格无法识别的地方。&lt;/p&gt; &lt;h2&gt;手动门禁系统或可保护工程技术&lt;/h2&gt; &lt;p&gt;将航空业的经验教训和核武器的本能放在一起比较，它们都指向了我们在人工智能增强型工作中需要的一种设计模式：刻意的“手动关卡”。&lt;/p&gt; &lt;p&gt;手动操作环节是指工作流程中由人接管控制权的某个点，这并非因为这是完成任务最快的方式，也不仅仅是为了安全联锁，而是为了锻炼和保持一项否则会退化的技能。其显著特点在于，它是人为选择的。在设计过程中，你需要决定你的组织必须保留哪些员工的技能，因为在关键时刻，这些技能至关重要。然后，你需要精心设计必要的机制来保持这些技能的熟练度。&lt;/p&gt; &lt;p&gt;想象一下，在一个主要依赖人工智能编写代码的软件团队中，这种机制会如何运作。团队会为他们最不能失去的技能——  &lt;a href="https://spectrum.ieee.org/tag/debugging" target="_blank"&gt;调试&lt;/a&gt;——设置一道人工屏障。当关键模块中出现缺陷时，指定的工程师（通常是初级工程师）必须首先重现故障，追溯根本原因，并编写一个能够捕获该缺陷的自动化测试，所有这些操作都必须在人工智能助手关闭的情况下进行。只有在工程师确定诊断结果后，模型才会重新启动，提出修复方案，生成替代方案，并在代码库中搜索类似的缺陷。然后，工程师会将自己的诊断结果与模型的分析结果进行比较。如果两者不一致，那就说明这种设计是有效的，它能够在问题发生之前就发现分歧，而不是在问题发生时才发现。&lt;/p&gt; &lt;p&gt;这种方法彻底改变了初级工程师的角色。如今人们的本能是让人工智能来完成入门级工作，因为它速度更快、成本更低。但有些工作并非可以随意削减的额外开销。它们是培养未来高级员工的基石，你应该像保护其他关键基础设施一样保护它们。或许它在本季度效率不高，但悄悄地拆除它，就等于把未来十年的能力抵押出去。&lt;/p&gt; &lt;h2&gt;为什么企业必须持续培养初级工程师&lt;/h2&gt; &lt;p&gt;这一切都不是免费的，假装免费是对那些必须签署预算的人的侮辱。人为设置的人工关卡，从本质上来说，短期内效率低于完全自动化。让初级员工从事基础工作并执行手动流程，现在付出一些成本，是为了将来能够获得更好的结果。&lt;/p&gt; &lt;p&gt;在如今这个以季度业绩评判大多数领导者的市场中，这种做法很难奏效。如果一位新聘高管手下有“不必要”的员工，而这些员工本可以被人工智能取代，那么在获得回报之前，董事会就会对此提出批评。只有那些能够免受这种压力影响的人才能从中获益：例如拥有控制权的创始人、私营公司、真正具有长远眼光的机构，或者像飞行员那样，愿意要求员工定期展示技能的监管机构。这意味着，最有可能保留自身专业技能的组织，是那些在结构上能够将短期利润投入到长期能力建设中的组织；其他所有组织都需要外部推动。&lt;/p&gt; &lt;p&gt;所以，论点可以用一句话概括：刻意降低效率并非浪费。在安全至关重要的工程领域，我们一直将其视为一种保障，而且我们是有意为之。随着人工智能接管那些需要专业知识才能完成的工作，明智的做法不是抵制自动化，而是从设计之初就将控制权牢牢掌握在手中——这样，即使自动化最终总会失效，也仍然有人能够胜任。&lt;/p&gt; &lt;div&gt;  &lt;br /&gt;&lt;/div&gt; &lt;div&gt;&lt;/div&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63278-%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD-%E6%8F%90%E5%8D%87-%E5%A4%B1%E5%8E%BB</guid>
      <pubDate>Fri, 04 Sep 2026 09:00:53 CST</pubDate>
    </item>
    <item>
      <title>Chronos-2 Small (112M)一个时间序列预测的强大模型</title>
      <link>https://itindex.net/detail/63277-chronos-small-112m</link>
      <description>&lt;p&gt;Chronos-2 是一个用于时间序列预测的基础模型，基于   &lt;a href="https://arxiv.org/abs/2403.07815" rel="nofollow"&gt;Chronos&lt;/a&gt; 和   &lt;a href="https://aws.amazon.com/blogs/machine-learning/fast-and-accurate-zero-shot-forecasting-with-chronos-bolt-and-autogluon/" rel="nofollow"&gt;Chronos-Bolt&lt;/a&gt; 构建。它提供了显著的能力改进，可以处理早期模型不支持的多种预测场景。它提供了显著的能力改进，可以处理早期模型不支持的多种预测场景。支持的场景包括（1）单变量时间序列预测：经典的时间序列预测任务；（2）跨项目学习：利用多个相关时间序列的信息进行预；（3）多变量预测：同时预测多个相关目标变量（3）带协变量的预测：仅过去协变量：在预测时已知的历史协变量（如过去的天气数据） + 已知未来协变量：预测期间已知的协变量（如节假日、计划事件）&lt;/p&gt; &lt;div&gt;  &lt;h2&gt;功能对比&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%8A%9F%E8%83%BD%E5%AF%B9%E6%AF%94"&gt;&lt;/a&gt;&lt;/div&gt; &lt;table&gt;  &lt;tr&gt;   &lt;th&gt;功能&lt;/th&gt;   &lt;th&gt;Chronos&lt;/th&gt;   &lt;th&gt;Chronos-Bolt&lt;/th&gt;   &lt;th&gt;Chronos-2&lt;/th&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;单变量预测&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;跨项目学习&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;多变量预测&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;仅过去协变量（实数/分类）&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;❌&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;已知未来协变量（实数/分类）&lt;/td&gt;   &lt;td&gt;🧩&lt;/td&gt;   &lt;td&gt;🧩&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;微调支持&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;   &lt;td&gt;✅&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;最大上下文长度&lt;/td&gt;   &lt;td&gt;512&lt;/td&gt;   &lt;td&gt;2048&lt;/td&gt;   &lt;td&gt;8192&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt; &lt;blockquote&gt;  &lt;p&gt;🧩 Chronos/Chronos-Bolt 不原生支持未来协变量，但可以与外部协变量回归器结合使用（参见    &lt;a href="https://auto.gluon.ai/stable/tutorials/timeseries/forecasting-chronos.html#incorporating-the-covariates" rel="nofollow"&gt;AutoGluon 教程&lt;/a&gt;）。这仅对每个时间步的效果进行建模，而不是跨时间的效果。相比之下，Chronos-2 原生支持所有协变量类型。&lt;/p&gt;&lt;/blockquote&gt; &lt;p&gt;更多关于 Chronos-2 的详细信息，请参阅   &lt;a href="https://www.arxiv.org/abs/2510.15821" rel="nofollow"&gt;技术报告&lt;/a&gt;。&lt;/p&gt; &lt;div&gt;  &lt;h2&gt;安装&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%AE%89%E8%A3%85"&gt;&lt;/a&gt;&lt;/div&gt; &lt;div&gt;  &lt;pre&gt;pip install &amp;apos;chronos-forecasting[extras]&amp;gt;=2.2&amp;apos; &amp;apos;matplotlib&amp;apos;&lt;/pre&gt;  &lt;div&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;h2&gt;快速开始&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%BF%AB%E9%80%9F%E5%BC%80%E5%A7%8B"&gt;&lt;/a&gt;&lt;/div&gt; &lt;div&gt;  &lt;h3&gt;加载模型&lt;/h3&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%8A%A0%E8%BD%BD%E6%A8%A1%E5%9E%8B"&gt;&lt;/a&gt;&lt;/div&gt; &lt;div&gt;  &lt;pre&gt;import os

# 使用仅 1 个 GPU（如果可用）
os.environ[&amp;quot;CUDA_VISIBLE_DEVICES&amp;quot;] = &amp;quot;0&amp;quot;

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from chronos import BaseChronosPipeline, Chronos2Pipeline

# 加载 Chronos-2 管道
# 推荐使用 GPU 加速推理，CPU 也支持（使用 device_map=&amp;quot;cpu&amp;quot;）
pipeline: Chronos2Pipeline = BaseChronosPipeline.from_pretrained(
    &amp;quot;amazon/chronos-2&amp;quot;, 
    device_map=&amp;quot;cuda&amp;quot;
)&lt;/pre&gt;  &lt;div&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;h3&gt;单变量预测&lt;/h3&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%8D%95%E5%8F%98%E9%87%8F%E9%A2%84%E6%B5%8B"&gt;&lt;/a&gt;&lt;/div&gt; &lt;div&gt;  &lt;pre&gt;# 加载数据（长格式 pandas DataFrame）
context_df = pd.read_csv(&amp;quot;https://autogluon.s3.amazonaws.com/datasets/timeseries/m4_hourly/train.csv&amp;quot;)
print(&amp;quot;输入数据框形状:&amp;quot;, context_df.shape)
display(context_df.head())

# 执行预测
pred_df = pipeline.predict_df(
    context_df, 
    prediction_length=24, 
    quantile_levels=[0.1, 0.5, 0.9]
)

print(&amp;quot;输出数据框形状:&amp;quot;, pred_df.shape)
display(pred_df.head())&lt;/pre&gt;  &lt;div&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;h3&gt;predict_df 参数说明&lt;/h3&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#predict_df-%E5%8F%82%E6%95%B0%E8%AF%B4%E6%98%8E"&gt;&lt;/a&gt;&lt;/div&gt; &lt;table&gt;  &lt;tr&gt;   &lt;th&gt;参数&lt;/th&gt;   &lt;th&gt;说明&lt;/th&gt;   &lt;th&gt;默认值&lt;/th&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;df&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;长格式 DataFrame，包含 id、timestamp 和目标列&lt;/td&gt;   &lt;td&gt;-&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;future_df&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;可选的未来协变量 DataFrame（同时存在于 df 和 future_df 中的列被视为已知未来协变量）&lt;/td&gt;   &lt;td&gt;None&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;id_column&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;时间序列标识符列&lt;/td&gt;   &lt;td&gt;&amp;quot;item_id&amp;quot;&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;timestamp_column&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;时间戳列&lt;/td&gt;   &lt;td&gt;&amp;quot;timestamp&amp;quot;&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;target&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;要预测的目标列名&lt;/td&gt;   &lt;td&gt;&amp;quot;target&amp;quot;&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;prediction_length&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;预测步数&lt;/td&gt;   &lt;td&gt;-&lt;/td&gt;&lt;/tr&gt;  &lt;tr&gt;   &lt;td&gt;    &lt;code&gt;quantile_levels&lt;/code&gt;&lt;/td&gt;   &lt;td&gt;要计算的分位数&lt;/td&gt;   &lt;td&gt;[0.1, 0.2, ..., 0.9]&lt;/td&gt;&lt;/tr&gt;&lt;/table&gt; &lt;p&gt;返回包含预测值（包括点预测和分位数）的 DataFrame。&lt;/p&gt; &lt;div&gt;  &lt;h2&gt;带协变量的预测&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%B8%A6%E5%8D%8F%E5%8F%98%E9%87%8F%E7%9A%84%E9%A2%84%E6%B5%8B"&gt;&lt;/a&gt;&lt;/div&gt; &lt;p&gt;Chronos-2 可以利用协变量来提高预测准确性。以下是一个实际的能源价格预测示例：&lt;/p&gt; &lt;div&gt;  &lt;pre&gt;# 能源价格预测配置
target = &amp;quot;target&amp;quot;  # 包含要预测的值的列名（能源价格）
prediction_length = 24  # 要预测的小时数
id_column = &amp;quot;id&amp;quot;  # 标识不同时间序列的列（国家/地区）
timestamp_column = &amp;quot;timestamp&amp;quot;  # 包含日期时间信息的列
timeseries_id = &amp;quot;DE&amp;quot;  # 要可视化的特定时间序列（德国）

# 加载历史能源价格和协变量的过去值
energy_context_df = pd.read_parquet(
    &amp;quot;https://autogluon.s3.amazonaws.com/datasets/timeseries/electricity_price/train.parquet&amp;quot;
)
energy_context_df[timestamp_column] = pd.to_datetime(energy_context_df[timestamp_column])

# 加载协变量的未来值
energy_test_df = pd.read_parquet(
    &amp;quot;https://autogluon.s3.amazonaws.com/datasets/timeseries/electricity_price/test.parquet&amp;quot;
)
energy_test_df[timestamp_column] = pd.to_datetime(energy_test_df[timestamp_column])
energy_future_df = energy_test_df.drop(columns=target)

# 使用协变量进行预测
pred_df = pipeline.predict_df(
    df=energy_context_df,
    future_df=energy_future_df,
    id_column=id_column,
    timestamp_column=timestamp_column,
    target=target,
    prediction_length=prediction_length,
    quantile_levels=[0.1, 0.5, 0.9],
)&lt;/pre&gt;  &lt;div&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;h2&gt;可视化预测结果&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%8F%AF%E8%A7%86%E5%8C%96%E9%A2%84%E6%B5%8B%E7%BB%93%E6%9E%9C"&gt;&lt;/a&gt;&lt;/div&gt; &lt;div&gt;  &lt;pre&gt;def plot_forecast(
    context_df,
    test_df,
    pred_df,
    timeseries_id,
    target,
    timestamp_column=&amp;quot;timestamp&amp;quot;,
    id_column=&amp;quot;item_id&amp;quot;,
):
    fig, ax = plt.subplots(figsize=(12, 3))
    
    # 筛选特定时间序列的数据
    ts_context = context_df[context_df[id_column] == timeseries_id].set_index(timestamp_column)[target]
    ts_ground_truth = test_df[test_df[id_column] == timeseries_id].set_index(timestamp_column)[target]
    ts_pred = pred_df[pred_df[id_column] == timeseries_id].set_index(timestamp_column)
    
    # 绘制历史数据、真实值和预测值
    ts_context.plot(ax=ax, label=f&amp;quot;历史 {target}&amp;quot;, color=&amp;quot;xkcd:azure&amp;quot;)
    ts_ground_truth.plot(ax=ax, label=f&amp;quot;未来 {target} (真实值)&amp;quot;, color=&amp;quot;xkcd:grass green&amp;quot;)
    ts_pred[&amp;quot;predictions&amp;quot;].plot(ax=ax, label=&amp;quot;预测&amp;quot;, color=&amp;quot;xkcd:violet&amp;quot;)
    
    # 绘制预测区间
    ax.fill_between(
        ts_pred.index,
        ts_pred[&amp;quot;0.1&amp;quot;],
        ts_pred[&amp;quot;0.9&amp;quot;],
        alpha=0.7,
        label=&amp;quot;预测区间&amp;quot;,
        color=&amp;quot;xkcd:light lavender&amp;quot;,
    )
    
    ax.axvline(x=ts_context.index[-1], color=&amp;quot;black&amp;quot;, linestyle=&amp;quot;--&amp;quot;, alpha=0.5)
    ax.legend(loc=&amp;quot;upper left&amp;quot;)
    ax.set_title(f&amp;quot;{target} 预测 - {timeseries_id}&amp;quot;)
    fig.show()&lt;/pre&gt;  &lt;div&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;h2&gt;支持的场景&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E6%94%AF%E6%8C%81%E7%9A%84%E5%9C%BA%E6%99%AF"&gt;&lt;/a&gt;&lt;/div&gt; &lt;ol&gt;  &lt;li&gt;单变量时间序列预测：经典的时间序列预测任务&lt;/li&gt;  &lt;li&gt;跨项目学习：利用多个相关时间序列的信息进行预测&lt;/li&gt;  &lt;li&gt;多变量预测：同时预测多个相关目标变量&lt;/li&gt;  &lt;li&gt;带协变量的预测：   &lt;ul&gt;    &lt;li&gt;仅过去协变量：在预测时已知的历史协变量（如过去的天气数据）&lt;/li&gt;    &lt;li&gt;已知未来协变量：预测期间已知的协变量（如节假日、计划事件）&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;&lt;/ol&gt; &lt;div&gt;  &lt;h2&gt;硬件要求&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E7%A1%AC%E4%BB%B6%E8%A6%81%E6%B1%82"&gt;&lt;/a&gt;&lt;/div&gt; &lt;ul&gt;  &lt;li&gt;GPU（推荐）：支持 CUDA 的 NVIDIA GPU，可显著加速推理&lt;/li&gt;  &lt;li&gt;CPU（支持）：可在 CPU 上运行，但推理速度较慢&lt;/li&gt;&lt;/ul&gt; &lt;div&gt;  &lt;h2&gt;参考&lt;/h2&gt;  &lt;a href="https://github.com/SuleynanAuir/Amazon-TSChronos#%E5%8F%82%E8%80%83"&gt;&lt;/a&gt;&lt;/div&gt; &lt;ul&gt;  &lt;li&gt;   &lt;a href="https://arxiv.org/abs/2403.07815" rel="nofollow"&gt;Chronos 论文&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://www.arxiv.org/abs/2510.15821" rel="nofollow"&gt;Chronos-2 技术报告&lt;/a&gt;&lt;/li&gt;  &lt;li&gt;   &lt;a href="https://auto.gluon.ai/stable/tutorials/timeseries/forecasting-chronos.html" rel="nofollow"&gt;AutoGluon 时间序列教程&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63277-chronos-small-112m</guid>
      <pubDate>Tue, 01 Sep 2026 10:30:03 CST</pubDate>
    </item>
    <item>
      <title>时间序列基础模型Moirai 2.0</title>
      <link>https://itindex.net/detail/63276-%E6%97%B6%E9%97%B4%E5%BA%8F%E5%88%97-%E5%9F%BA%E7%A1%80-%E6%A8%A1%E5%9E%8B</link>
      <description>&lt;p&gt;时间序列基础模型是过去两年最热的方向之一：Google 的   &lt;a href="https://www.biaodianfu.com/timesfm/"&gt;TimesFM&lt;/a&gt;、亚马逊的  &lt;a href="https://www.biaodianfu.com/chronos-2/"&gt; Chronos&lt;/a&gt;、Datadog 的 TOTO 先后登场，都想做一个模型预测所有领域。2025 年 8 月，Salesforce 发布的 Moirai 2.0 给出了一个反直觉的答案——把架构做得更简单，模型做得更小，反而更强，论文标题直接就叫《When Less Is More》（少即是多）。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="479" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/Moirai.png" width="897"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;Moirai 2.0 是一个纯解码器（decoder-only）时间序列基础模型，在 3600 万条序列、约 2950 亿观测值的语料上预训练，直接输出 9 个分位数（0.1～0.9）的概率预测。它的 Small 版本只有 11.4M 参数，却比上一代 311M 参数的 Moirai 1.0-Large 更准，推理快约 2 倍、体积小约 30 倍。&lt;/p&gt;
 &lt;h2&gt;为什么需要 Moirai 2.0？&lt;/h2&gt;
 &lt;p&gt;先给结论：Moirai 2.0 的诞生，是对时序基础模型就该又大又复杂这一假设的系统性纠偏——它砍掉了前代三个最复杂的组件，换来更准、更快、更小。&lt;/p&gt;
 &lt;p&gt;时间序列本身很难：非平稳、多尺度、采样不规则、观测有缺失，跨领域泛化极难。传统做法是每个场景训一个专用模型（ETS、ARIMA，或 PatchTST、N-BEATS 这类深度模型），维护成本高。基础模型的思路是像 GPT 之于文本一样，用海量跨域数据预训练一个通用模型，新场景零样本直接用。Moirai 1.0（2024 年，ICML Oral）正是这条路线最早的践行者之一：基于 LOTSA 数据集（27B 观测、9 大领域）预训练，曾登顶 GIFT-Eval 排行榜。&lt;/p&gt;
 &lt;p&gt;但 1.0 的成功也暴露了三个结构性短板：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;数据利用率低：掩码编码器架构下，每条训练样本只产生一个损失，全序列仅约 15% 的 token 真正参与了损失计算，大部分训练信号被浪费。&lt;/li&gt;
  &lt;li&gt;多补丁设计过重：为适配不同时间频率引入了多补丁输入，反而限制了跨频率学习，还增加了计算开销。&lt;/li&gt;
  &lt;li&gt;混合分布输出难优化：用混合分布（mixture of distributions）做概率预测，理论上灵活，实践中梯度不稳定、优化复杂，收益却不明显。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;Moirai 2.0 的回应是三个砍：掩码编码器 → 纯解码器；多补丁 → 单补丁；混合分布 → 分位数。结果是一条样本产生 T−1 个训练损失、训练和推理都大幅简化、概率输出天然可用——用论文里的原话，这套改动让它比自己家族里更大的模型表现更好。&lt;/p&gt;
 &lt;p&gt;这条演进路线的时间线如下：&lt;/p&gt;
 &lt;table width="760"&gt;

  &lt;tr&gt;
   &lt;td&gt;时间&lt;/td&gt;
   &lt;td&gt;事件&lt;/td&gt;
   &lt;td&gt;意义&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.02&lt;/td&gt;
   &lt;td&gt;Moirai 1.0 论文发布（arXiv:2402.02592）&lt;/td&gt;
   &lt;td&gt;最早的大规模通用时序基础模型之一&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.03&lt;/td&gt;
   &lt;td&gt;Uni2TS 框架与 LOTSA 数据开源&lt;/td&gt;
   &lt;td&gt;代码、数据、权重全部开放&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.05&lt;/td&gt;
   &lt;td&gt;论文被 ICML 2024 接收（Oral）&lt;/td&gt;
   &lt;td&gt;学术界的标志性认可&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.06&lt;/td&gt;
   &lt;td&gt;Moirai 1.1 发布&lt;/td&gt;
   &lt;td&gt;低频数据（年度/季度）NMAE 提升约 20%&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.10&lt;/td&gt;
   &lt;td&gt;Moirai-MoE 发布&lt;/td&gt;
   &lt;td&gt;混合专家架构版本&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2024.11&lt;/td&gt;
   &lt;td&gt;GIFT-Eval 基准发布&lt;/td&gt;
   &lt;td&gt;时序基础模型的公共评测场，Moirai-Large 曾登顶&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2025.08&lt;/td&gt;
   &lt;td&gt;Moirai 2.0 发布（Salesforce 官方博客）&lt;/td&gt;
   &lt;td&gt;架构全面转向 decoder-only，非数据泄露模型中 MASE 第一&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;2025.11&lt;/td&gt;
   &lt;td&gt;Moirai 2.0 论文发表（arXiv:2511.11698）&lt;/td&gt;
   &lt;td&gt;完整公开架构、数据与消融实验&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;h2&gt;五大核心设计：Less 究竟砍掉了什么&lt;/h2&gt;
 &lt;p&gt;先给结论：Moirai 2.0 的少，落在架构、输出、解码、训练四个层面，每一刀都精准对准 1.0 的一个痛点。下图是它的端到端流水线。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="681" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/less.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;h3&gt;纯解码器架构：让每个 token 都「练到」&lt;/h3&gt;
 &lt;p&gt;如果说掩码编码器像老师拿着全文让学生做一道阅读理解题，那么纯解码器就像顺着文章往下读，每读一个词都预测下一个词——每个位置都参与训练，预测永远只依赖当前及之前的内容（因果性）。&lt;/p&gt;
 &lt;p&gt;这一改动带来了三方面收益。其一，训练信号量级提升：一条长度为 T 的序列，decoder-only 能产生 T−1 个损失，而掩码编码器只有 1 个，数据利用率从约 15% 的 token 参与损失提升到全部 token。其二，推理天然自回归：预测未来就是接着往下写，与模型的训练目标完全一致。其三，可复用 KV 缓存：重复/扩展预测时，已算过的键值对直接缓存复用，论文报告迭代式扩展预测最多可提速 17 倍。&lt;/p&gt;
 &lt;p&gt;架构上它与现代 LLM 同构：输入补丁经残差块（SiLU 激活）投影为 token，送入多层因果多头自注意力 + 前馈网络的堆叠，最后输出投影把每个 token 映射为多令牌、多分位数预测。small / base / large 三个变体分别是 12 / 24 / 48 层。&lt;/p&gt;
 &lt;h3&gt;分位数预测与 Pinball 损失：原生概率输出&lt;/h3&gt;
 &lt;p&gt;先给结论：Moirai 2.0 不预测未来是多少，而是预测未来有 90% 概率落在哪个区间，这个区间由 9 个分位数（q=0.1, 0.2, …, 0.9）直接给出，无需事后拟合分布。&lt;/p&gt;
 &lt;p&gt;分位数是什么？把历史数据按大小排序，第 q 分位数就是有 q 比例的数据不超过它的那个值。预测时，模型对每个未来时间步直接输出 9 个分位数值：q=0.1 和 q=0.9 之间的区间就是 80% 置信带，q=0.5（中位数）则天然可作为点预测。如下图所示。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="524" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/Pinball.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;训练时用分位数损失（又称 Pinball 损失）直接优化这 9 个分位数。对时间步 t 的真实值 $y_t$ 与预测分位数值 $\hat{y}_t^{(q)}$：&lt;/p&gt;
 &lt;p&gt;$$\ell_q(y_t, \hat{y}_t^{(q)}) = \begin{cases} q\,(y_t – \hat{y}_t^{(q)}) &amp;amp; \text{if } y_t \ge \hat{y}_t^{(q)} \\ (1-q)\,(\hat{y}_t^{(q)} – y_t) &amp;amp; \text{if } y_t &amp;lt; \hat{y}_t^{(q)} \end{cases}$$&lt;/p&gt;
 &lt;p&gt;逐项拆解这个公式：当真实值高于预测值（$y_t \ge \hat{y}_t^{(q)}$，低估了），惩罚是 $q \times$ 偏差；当真实值低于预测值（高估了），惩罚是 $(1-q) \times$ 偏差。系数 q 的作用是不对称：对高分位数（如 q=0.9），低估（真实值比预测高）会吃到 0.9 倍偏差的惩罚，比高估（0.1 倍）重得多——这迫使模型把 q=0.9 这条线抬高，保证90% 的值都不超过它，从而形成正确的分位数排序与间距。总损失是所有分位数、所有时间步的平均：&lt;/p&gt;
 &lt;p&gt;$$\mathcal{L}_Q = \frac{1}{H|Q|} \sum_{t=1}^{H} \sum_{q \in Q} \ell_q\big(y_t, \hat{y}_t^{(q)}\big), \quad Q = \{0.1, 0.2, \dots, 0.9\}$$&lt;/p&gt;
 &lt;p&gt;为什么直接回归分位数比混合分布更好？因为分位数损失是概率预测评估指标 CRPS 的离散近似——训练目标和评测指标对齐了，且逐点损失对异常值稳健、梯度稳定，没有混合分布那样的模式坍缩与优化困难。消融实验显示，仅这一项改动就把 MASE 从 0.850 拉到 0.744，是整个模型中收益最大的一刀。&lt;/p&gt;
 &lt;p&gt;这套设计的局限与应对：默认 9 个分位数等权，对重尾分布表达有限；但损失函数支持对不同分位数加权（如容量规划强调 q=0.9 的高分位），可按业务定制。最佳场景是金融风控（要上下界）、运维容量规划（要 P95/P99）、任何预测错了要付出不对称代价的场景。&lt;/p&gt;
 &lt;h3&gt;多令牌预测：一次吐出一串未来&lt;/h3&gt;
 &lt;p&gt;如果说单令牌预测是一个词一个词往外蹦，那么多令牌预测就是一次说出一句完整的话。Moirai 2.0 的每个输出 token 同时预测 $n_{token}$ 个未来补丁（patch），输出投影从 $\mathbb{R}^d$ 映射到 $\mathbb{R}^{n_{token} \times n_q \times p}$，一次前向就覆盖一段更长的预测范围。&lt;/p&gt;
 &lt;p&gt;这么做的直接收益是效率与稳定性：自回归迭代次数成倍减少，长预测范围下的误差累积也随之降低——论文的消融实验证实，多令牌预测在最终版本中贡献了约 0.011 的 MASE 改善（0.739 → 0.728 的路径上）。&lt;/p&gt;
 &lt;h3&gt;自回归多分位数解码：先展开、再折叠&lt;/h3&gt;
 &lt;p&gt;这里有个必须解决的矛盾：每一步输出 9 个分位数，自回归时该把哪一个当作输入传给下一步？取中位数会丢掉全部不确定性信息；直接传 9 个又维度不匹配。Moirai 2.0 的解法是一个深度为 2 的扩展—折叠（expand-then-fold）束搜索：&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="584" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/zihuigui.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;具体到每一步：第一步直接用上下文预测出 9 个分位数；展开阶段，把上一步的每个分位数值当作输入分别解码，得到 $9 \times 9 = 81$ 个候选；折叠阶段，对这 81 个候选在每一步取对应分位数（$\hat{y}_{t+1}^{(q)} \leftarrow \text{Quantile}_q(\text{candidates})$），聚合回 9 个标准分位数；如此循环直到预测范围结束。这相当于在保留全概率信息的同时，把多值自回归变成了可行计算。&lt;/p&gt;
 &lt;h3&gt;训练三件套：防泄漏、抗缺失、滤噪声&lt;/h3&gt;
 &lt;p&gt;除了架构，训练策略还有三个看似细节、实则关键的机制：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;防泄漏归一化：decoder-only 模型若对整个序列做全局实例归一化，归一化统计量会偷看未来的 70% 数据。因此只用序列前 30% 计算均值和方差，后 70% 用于因果预训练，杜绝未来信息泄露。&lt;/li&gt;
  &lt;li&gt;Patch 级随机掩码：训练时随机掩掉 50% 的输入补丁，让模型学会从残缺输入中预测，增强对真实场景缺失值的鲁棒性。消融实验表明它单独使用反而降精度，但与其他策略组合时有正向贡献。&lt;/li&gt;
  &lt;li&gt;数据过滤：既然归一化只看前 30%，后 70% 段可能发生分布偏移。用 Z-score 异常检测对比两段统计量，把不可预测的低质量序列从预训练语料中滤掉，稳定收敛。&lt;/li&gt;
&lt;/ul&gt;
 &lt;h2&gt;预训练数据：3600 万条序列从哪来&lt;/h2&gt;
 &lt;p&gt;先给结论：Moirai 2.0 在 3600 万条序列、约 2950 亿观测值的混合语料上预训练，真实数据与合成数据并重，覆盖运维、能源、医疗、经济、交通、电商等主要领域——这是它零样本跨域能力的根基。&lt;/p&gt;
 &lt;table width="760"&gt;

  &lt;tr&gt;
   &lt;td&gt;数据来源&lt;/td&gt;
   &lt;td&gt;序列数&lt;/td&gt;
   &lt;td&gt;观测数&lt;/td&gt;
   &lt;td&gt;说明&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;GIFT-Eval Pretrain（无泄露版）&lt;/td&gt;
   &lt;td&gt;约 325 万&lt;/td&gt;
   &lt;td&gt;约 2300 亿&lt;/td&gt;
   &lt;td&gt;LOTSA 的子集，精心挑选避免与评测任务重叠&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;GIFT-Eval TrainTest 训练集&lt;/td&gt;
   &lt;td&gt;约 14.4 万&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;评测集训练部分，用于增强&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Chronos-Mixup（合成）&lt;/td&gt;
   &lt;td&gt;3000 万&lt;/td&gt;
   &lt;td&gt;约 630 亿&lt;/td&gt;
   &lt;td&gt;基于 TSMixup 生成，仅用 Chronos 非泄露子集&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;KernelSynth（合成）&lt;/td&gt;
   &lt;td&gt;100 万&lt;/td&gt;
   &lt;td&gt;约 10.2 亿&lt;/td&gt;
   &lt;td&gt;高斯过程合成：趋势、局部变化、季节核随机组合&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Salesforce 内部遥测（匿名）&lt;/td&gt;
   &lt;td&gt;约 215 万&lt;/td&gt;
   &lt;td&gt;约 14.8 亿&lt;/td&gt;
   &lt;td&gt;日粒度云监控数据，覆盖约一年（自 2024 年 1 月起）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;合计&lt;/td&gt;
   &lt;td&gt;3600 万&lt;/td&gt;
   &lt;td&gt;约 2950 亿&lt;/td&gt;
   &lt;td&gt;真实 + 合成混合，8 大领域&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;值得注意两点。第一，合成数据占比很高（Chronos-Mixup 3000 万条占了近九成），说明在真实高质量时序数据稀缺的情况下，合成数据是扩大覆盖面的有效手段——这一点与 Chronos 论文的做法一脉相承。第二，语料的领域分布并不均衡：自然/环境类序列偏少，这直接导致了后面实验里Natural 领域表现偏弱的结果。&lt;/p&gt;
 &lt;h2&gt;实验结果：小模型如何打赢大模型&lt;/h2&gt;
 &lt;p&gt;先给结论：在 GIFT-Eval 基准（55 个数据集、97 种任务配置、37 个模型参赛）上，参数最少的 Moirai 2.0-Small（11.4M）反而在所有尺寸中表现最好，并在非数据泄露模型中拿到 MASE 第一。这是一份缩小模型反而变强的实证。&lt;/p&gt;
 &lt;h3&gt;缩放实验：参数不是越多越好&lt;/h3&gt;
 &lt;table width="380"&gt;

  &lt;tr&gt;
   &lt;td width="140"&gt;模型&lt;/td&gt;
   &lt;td width="90"&gt;参数（M）&lt;/td&gt;
   &lt;td width="75"&gt;MASE ↓&lt;/td&gt;
   &lt;td width="75"&gt;CRPS ↓&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="140"&gt;Moirai 2.0 Small&lt;/td&gt;
   &lt;td width="90"&gt;11.4&lt;/td&gt;
   &lt;td width="75"&gt;0.728&lt;/td&gt;
   &lt;td width="75"&gt;0.516&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="140"&gt;Moirai 2.0 Base&lt;/td&gt;
   &lt;td width="90"&gt;87.1&lt;/td&gt;
   &lt;td width="75"&gt;0.732&lt;/td&gt;
   &lt;td width="75"&gt;0.525&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="140"&gt;Moirai 2.0 Large&lt;/td&gt;
   &lt;td width="90"&gt;305&lt;/td&gt;
   &lt;td width="75"&gt;0.743&lt;/td&gt;
   &lt;td width="75"&gt;0.530&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;从 11.4M 扩到 305M，性能不升反降。论文给出的解释是：当前预训练数据的规模与多样性和模型容量不匹配，仅堆参数无效——这打破了参数量 = 精度的直觉，也把未来方向指向了数据扩展而非模型扩展。&lt;/p&gt;
 &lt;h3&gt;消融实验：哪一刀收益最大&lt;/h3&gt;
 &lt;table width="788"&gt;

  &lt;tr&gt;
   &lt;td&gt;变体&lt;/td&gt;
   &lt;td&gt;预训练数据&lt;/td&gt;
   &lt;td width="68"&gt;损失&lt;/td&gt;
   &lt;td width="84"&gt;架构&lt;/td&gt;
   &lt;td&gt;投影&lt;/td&gt;
   &lt;td&gt;多令牌&lt;/td&gt;
   &lt;td&gt;递归解码&lt;/td&gt;
   &lt;td&gt;随机掩码&lt;/td&gt;
   &lt;td&gt;MASE ↓&lt;/td&gt;
   &lt;td&gt;CRPS ↓&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Moirai 1.0 small&lt;/td&gt;
   &lt;td&gt;GIFT-Eval Pretrain&lt;/td&gt;
   &lt;td width="68"&gt;分布&lt;/td&gt;
   &lt;td width="84"&gt;enc-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;0.946&lt;/td&gt;
   &lt;td&gt;0.650&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v0&lt;/td&gt;
   &lt;td&gt;GIFT-Eval Pretrain&lt;/td&gt;
   &lt;td width="68"&gt;分布&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;0.929&lt;/td&gt;
   &lt;td&gt;0.647&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v1&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分布&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;0.850&lt;/td&gt;
   &lt;td&gt;0.580&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v2&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分位数&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;0.744&lt;/td&gt;
   &lt;td&gt;0.553&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v3&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分位数&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;0.736&lt;/td&gt;
   &lt;td&gt;0.533&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v4&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分位数&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;—&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;0.772&lt;/td&gt;
   &lt;td&gt;0.560&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;v5&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分位数&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;线性&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;0.739&lt;/td&gt;
   &lt;td&gt;0.527&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Moirai 2.0&lt;/td&gt;
   &lt;td&gt;新语料&lt;/td&gt;
   &lt;td width="68"&gt;分位数&lt;/td&gt;
   &lt;td width="84"&gt;dec-only&lt;/td&gt;
   &lt;td&gt;残差块&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;✓&lt;/td&gt;
   &lt;td&gt;0.728&lt;/td&gt;
   &lt;td&gt;0.516&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;这条消融链清晰展示了每个决策的贡献：decoder-only 骨架带来小幅提升（0.946 → 0.929）；新语料是第二大步（0.929 → 0.850）；分位数损失是单项收益最大的一刀（0.850 → 0.744）；递归分位数解码（v3）与多令牌预测（v5）各添一分；随机掩码单独用反而降（v4 的 0.772 高于 v3），但在完整组合里有利；最后把线性投影换成残差块，凑出最终成绩。&lt;/p&gt;
 &lt;h3&gt;与前代的全面对比&lt;/h3&gt;
 &lt;table width="548"&gt;

  &lt;tr&gt;
   &lt;td width="152"&gt;对比维度&lt;/td&gt;
   &lt;td width="173"&gt;Moirai 1.0&lt;/td&gt;
   &lt;td width="222"&gt;Moirai 2.0&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;架构&lt;/td&gt;
   &lt;td width="173"&gt;掩码编码器&lt;/td&gt;
   &lt;td width="222"&gt;纯解码器&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;输入补丁&lt;/td&gt;
   &lt;td width="173"&gt;多补丁（多频率）&lt;/td&gt;
   &lt;td width="222"&gt;单补丁&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;概率输出&lt;/td&gt;
   &lt;td width="173"&gt;混合分布 + 采样&lt;/td&gt;
   &lt;td width="222"&gt;9 个分位数直接输出&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;损失&lt;/td&gt;
   &lt;td width="173"&gt;分布 NLL&lt;/td&gt;
   &lt;td width="222"&gt;分位数（Pinball）损失&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;训练数据利用率&lt;/td&gt;
   &lt;td width="173"&gt;约 15% token 参与损失&lt;/td&gt;
   &lt;td width="222"&gt;全部 token，T−1 个损失/样本&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;KV 缓存推理&lt;/td&gt;
   &lt;td width="173"&gt;不支持&lt;/td&gt;
   &lt;td width="222"&gt;支持，重复预测最多提速 17 倍&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;模型规模（最优款）&lt;/td&gt;
   &lt;td width="173"&gt;Large 约 311M&lt;/td&gt;
   &lt;td width="222"&gt;Small 仅 11.4M（小约 30 倍）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;推理速度&lt;/td&gt;
   &lt;td width="173"&gt;基准&lt;/td&gt;
   &lt;td width="222"&gt;快约 2 倍&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="152"&gt;精度&lt;/td&gt;
   &lt;td width="173"&gt;基准&lt;/td&gt;
   &lt;td width="222"&gt;MASE +16%、CRPS +13%&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;领域层面的结论也值得关注：Moirai 2.0 在 Web/CloudOps、金融等领域表现突出（与其内部云监控数据有关），在自然/环境序列上明显偏弱（预训练语料该类数据太少），交通领域则略逊于 Moirai 1.0。按预测长度分组，短序列排名第 4、中期第 6、长期第 8——优势随预测范围拉长而收窄。&lt;/p&gt;
 &lt;p&gt;效率上它处于最优性价比区间：同批竞品中，Kairos-50M 更快但精度明显落后（参数还是它的近 5 倍），Granite-FlowState-R1 精度略高但推理慢约 3 倍。Moirai 2.0 在精度、速度、体积三者的权衡上最均衡。&lt;/p&gt;
 &lt;h2&gt;快速上手：五分钟跑通 Moirai 2.0&lt;/h2&gt;
 &lt;p&gt;先给结论：装一个 uni2ts 库，加载 Hugging Face 上的预训练权重，三行代码就能做零样本概率预测，Small 模型在普通 GPU 上即可跑。&lt;/p&gt;
 &lt;p&gt;安装（官方推荐的 PyPI 方式，框架基于 PyTorch，推理依赖 GluonTS）：&lt;/p&gt;
 &lt;pre&gt;pip install uni2ts&lt;/pre&gt;
 &lt;p&gt;零样本预测完整示例——从 pandas 宽表出发，预测并可视化：&lt;/p&gt;
 &lt;pre&gt;import pandas as pd
import matplotlib.pyplot as plt
from gluonts.dataset.pandas import PandasDataset
from gluonts.dataset.split import split
from uni2ts.eval_util.plot import plot_single
from uni2ts.model.moirai2 import Moirai2Forecast, Moirai2Module

# 1) 读入宽表：行 = 时间，列 = 各条序列
df = pd.read_csv(ts_wide.csv, index_col=0, parse_dates=True)

# 2) 转为 GluonTS 数据集并做滚动切分
ds = PandasDataset(dict(df))
train, test_template = split(ds, offset=-100)  # 末 100 步作为测试
test_data = test_template.generate_instances(
    prediction_length=100,  # 预测长度
    windows=1,              # 滚动窗口数
    distance=100,           # 窗口间隔
)

# 3) 加载预训练权重（Moirai 2.0 Small，仅 11.4M 参数）
model = Moirai2Forecast(
    module=Moirai2Module.from_pretrained(Salesforce/moirai-2.0-R-small),
    prediction_length=100,
    context_length=1680,
    target_dim=1,
    feat_dynamic_real_dim=0,
    past_feat_dynamic_real_dim=0,
)

# 4) 零样本预测 + 可视化
predictor = model.create_predictor(batch_size=32)
forecasts = predictor.predict(test_data.input)
inp = next(iter(test_data.input))
label = next(iter(test_data.label))
forecast = next(iter(forecasts))
plot_single(inp, label, forecast, context_length=200, name=pred, show_label=True)
plt.show()
&lt;/pre&gt;
 &lt;p&gt;关键参数速查：&lt;/p&gt;
 &lt;table width="760"&gt;

  &lt;tr&gt;
   &lt;td&gt;参数&lt;/td&gt;
   &lt;td&gt;含义&lt;/td&gt;
   &lt;td&gt;建议&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;MODEL = moirai2&lt;/td&gt;
   &lt;td&gt;选择模型家族&lt;/td&gt;
   &lt;td&gt;可选 moirai（1.x）/ moirai-moe / moirai2&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;SIZE = small&lt;/td&gt;
   &lt;td&gt;模型尺寸&lt;/td&gt;
   &lt;td&gt;small / base / large；论文结论 small 已最优&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;prediction_length&lt;/td&gt;
   &lt;td&gt;预测长度（步数）&lt;/td&gt;
   &lt;td&gt;任意正整数，按业务设定&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;context_length&lt;/td&gt;
   &lt;td&gt;输入上下文长度&lt;/td&gt;
   &lt;td&gt;建议 1000+（论文示例用 1680）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;batch_size&lt;/td&gt;
   &lt;td&gt;推理批大小&lt;/td&gt;
   &lt;td&gt;按显存调整，small 模型 32 起步&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;target_dim&lt;/td&gt;
   &lt;td&gt;目标变量维度&lt;/td&gt;
   &lt;td&gt;单变量为 1；多变量按列数设为 N&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;两个落地提示：一是多变量数据会被拆成多条独立单变量序列分别预测（模型本身不支持跨变量联合建模）；二是上下文给足、预测长度别一口气拉太长，长 horizon 的精度衰减是它的已知短板。&lt;/p&gt;
 &lt;h2&gt;应用场景与落地建议&lt;/h2&gt;
 &lt;p&gt;先给结论：Moirai 2.0 最适合多领域、多频率、需要概率边界、算力预算有限的预测任务；它是零样本快速起步的最佳选择之一，也是中小规模部署里性价比最高的时序基础模型。&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;IT 运维容量规划：对 CPU、内存、流量等云监控指标做多步预测。为什么适配——预训练语料里就有大量 Salesforce 内部云监控数据，Web/CloudOps 是它最强的领域；分位数输出直接给出 P90/P99 上界，正好服务按峰值扩容的决策；KV 缓存 + 11.4M 参数让它能在边缘节点低延迟运行。&lt;/li&gt;
  &lt;li&gt;金融风控与财务预测：预测波动率、交易量、营收指标并给出置信区间。为什么适配——分位数损失对齐 CRPS、对异常值稳健，天然适配尾部风险分析；金融领域的跨域表现位居前列。&lt;/li&gt;
  &lt;li&gt;零售与供应链需求预测：零样本预测新品或冷启动品类的销量。为什么适配——跨域泛化能力强，无需为每个 SKU 单独训练；合成数据训练的多样性让它对没见过的销售形态也有一定鲁棒性。&lt;/li&gt;
  &lt;li&gt;在线/边缘实时预测：小模型 + KV 缓存意味着毫秒级推理，适合设备端或网关侧的流式预测。为什么适配——模型体积小（Small 约 45MB 权重），可在资源受限环境部署。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;与同赛道模型的定位对比（帮你选型）：&lt;/p&gt;
 &lt;table width="760"&gt;

  &lt;tr&gt;
   &lt;td&gt;模型&lt;/td&gt;
   &lt;td&gt;架构&lt;/td&gt;
   &lt;td&gt;输出&lt;/td&gt;
   &lt;td&gt;模型规模&lt;/td&gt;
   &lt;td&gt;特点&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Moirai 2.0&lt;/td&gt;
   &lt;td&gt;decoder-only&lt;/td&gt;
   &lt;td&gt;9 个分位数&lt;/td&gt;
   &lt;td&gt;11.4M / 87M / 305M&lt;/td&gt;
   &lt;td&gt;效率与精度最均衡，概率输出开箱即用&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;TimesFM 系列&lt;/td&gt;
   &lt;td&gt;decoder-only&lt;/td&gt;
   &lt;td&gt;点预测为主&lt;/td&gt;
   &lt;td&gt;200M 级&lt;/td&gt;
   &lt;td&gt;Google 出品，生态成熟，点预测基线强&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Chronos / Chronos-Bolt&lt;/td&gt;
   &lt;td&gt;编码器-解码器&lt;/td&gt;
   &lt;td&gt;采样分布&lt;/td&gt;
   &lt;td&gt;20M～710M&lt;/td&gt;
   &lt;td&gt;Amazon 出品，Bolt 推理极快&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Moirai 1.x / Moirai-MoE&lt;/td&gt;
   &lt;td&gt;掩码编码器 / MoE&lt;/td&gt;
   &lt;td&gt;混合分布&lt;/td&gt;
   &lt;td&gt;14M～311M&lt;/td&gt;
   &lt;td&gt;前代家族，MoE 版本追求更大容量&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;落地时注意三点反模式：一是别把预测长度设成上下文的好几倍，长 horizon 衰减明显；二是需要协变量（天气、节假日）强介入的场景，先考虑微调或传统模型，它目前不消费外部特征；三是追求多变量联合概率（如多仓库存联动）的场景它做不到，只能逐条预测。需要针对性微调时，Uni2TS 自带 CLI：&lt;/p&gt;
 &lt;p&gt;python cli/train.py data=etth1 model=moirai_1.1_R_base trainer.max_epochs=50&lt;/p&gt;
 &lt;h2&gt;局限与破局：四条边界与各自的解法&lt;/h2&gt;
 &lt;p&gt;先给结论：Moirai 2.0 的边界清晰且诚实——它把能做和暂时不能做分得很开，每条局限也都有明确的破局路径。&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;只做单变量：多变量被拆成多条独立序列，丢失了变量间的联合依赖（如多条指标的相关性）。破局方向：论文明确把多变量联合概率预测列为后续工作，或先用合成数据训练多变量分支；短期可自行用多变量模型（如 DeepAR）做对照。&lt;/li&gt;
  &lt;li&gt;不消费外部协变量：天气、节假日、促销等外生特征无法输入，限制了强外部驱动场景（如零售大促）。破局方向：对协变量敏感的任务用传统统计模型或微调接入；社区也在探索把协变量编码进补丁嵌入的扩展。&lt;/li&gt;
  &lt;li&gt;长 horizon 精度衰减：短序列排名第 4、长期降到第 8，多令牌 + 迭代解码只是缓解而非根治。破局方向：论文建议长程建模专项研究 + 数据扩展；实践上分段滚动预测（每次预测一段，回填再推）比一口气预测到底更稳。&lt;/li&gt;
  &lt;li&gt;参数扩展收益为负：4M → 305M 性能反降，说明瓶颈在数据而非模型。破局方向：先扩大预训练语料规模与多样性（论文的明确未来工作），再谈增大模型；应用侧直接用 Small 即可，不必盲目上 Large。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;展望未来，论文给出的方向包括结合 LLM 推理能力的时序分析、文本/图像/时序的多模态基础模型，以及数据扩展与长程建模。对一个少即是多的模型来说，下一步的想象力在于：当数据这块短板补上之后，更小的架构还能撬动多大的性能天花板。&lt;/p&gt;
 &lt;div&gt;

  &lt;strong&gt;相关文章:&lt;/strong&gt;  &lt;ol&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/uplift-model/" rel="bookmark" title="&amp;#33829;&amp;#38144;&amp;#22686;&amp;#30410;&amp;#27169;&amp;#22411;(Uplift Model)"&gt;营销增益模型(Uplift Model)&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/chronos-2/" rel="bookmark" title="Amazon&amp;#26102;&amp;#24207;&amp;#39044;&amp;#27979;&amp;#27169;&amp;#22411;Chronos-2"&gt;Amazon时序预测模型Chronos-2&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/transformer/" rel="bookmark" title="&amp;#33258;&amp;#28982;&amp;#35821;&amp;#35328;&amp;#22788;&amp;#29702;&amp;#20043;Transformer"&gt;自然语言处理之Transformer&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>器→工具 工具软件 数据 术→技巧 时序预测</category>
      <guid isPermaLink="true">https://itindex.net/detail/63276-%E6%97%B6%E9%97%B4%E5%BA%8F%E5%88%97-%E5%9F%BA%E7%A1%80-%E6%A8%A1%E5%9E%8B</guid>
      <pubDate>Mon, 31 Aug 2026 22:01:50 CST</pubDate>
    </item>
    <item>
      <title>基于LLM的开源时序预测模型Lag-Llama</title>
      <link>https://itindex.net/detail/63275-llm-%E5%BC%80%E6%BA%90-%E9%A2%84%E6%B5%8B</link>
      <description>&lt;p&gt;Lag-Llama 是时间序列预测领域的第一个开源基础模型：它用 LLaMA 式的解码器 Transformer，把一段历史序列翻译成对未来概率分布的预测，无需针对每个数据集重新训练。2023 年 10 月由 Kashif Rasul 等 18 位作者发布（arXiv:2310.08278，代码 Apache-2.0 开源），2024 年 2 月官方权重在 Hugging Face 开放。它不解决多变量相关性这类大问题，只做一件事——单变量序列的通用概率预测，却因此成为时序基础模型赛道的起点：此后 TimesFM、Chronos、Moirai 相继登场，几乎都以它为对比基准。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="575" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/Lag-Llama.png" width="1020"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;table width="735"&gt;

  &lt;tr&gt;
   &lt;td width="98"&gt;时间&lt;/td&gt;
   &lt;td width="257"&gt;事件&lt;/td&gt;
   &lt;td&gt;意义&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="98"&gt;2017&lt;/td&gt;
   &lt;td width="257"&gt;DeepAR 发布&lt;/td&gt;
   &lt;td&gt;概率预测的经典深度模型，但每个数据集都要单独训练&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="98"&gt;2022–2023&lt;/td&gt;
   &lt;td width="257"&gt;PatchTST、OneFitsAll 等出现&lt;/td&gt;
   &lt;td&gt;探索通用架构：有人直接把 GPT-2 搬到时序上（OneFitsAll），效果平平&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="98"&gt;2023-10&lt;/td&gt;
   &lt;td width="257"&gt;Lag-Llama 论文发布&lt;/td&gt;
   &lt;td&gt;首个时序预测基础模型，架构、数据、权重全部开源&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="98"&gt;2024-02&lt;/td&gt;
   &lt;td width="257"&gt;官方权重 + Colab Demo 开放&lt;/td&gt;
   &lt;td&gt;零样本预测开箱即用，任何频率、任意预测长度&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="98"&gt;2024 全年&lt;/td&gt;
   &lt;td width="257"&gt;TimesFM、Chronos、Moirai 相继发布&lt;/td&gt;
   &lt;td&gt;时序基础模型赛道爆发，Lag-Llama 成为公认的早期基准&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;h2&gt;为什么需要时序预测基础模型&lt;/h2&gt;
 &lt;p&gt;时序预测长期处于碎片化状态：换一个领域、换一种采样频率，就要重新训练一个模型。DeepAR、N-BEATS、TFT 等模型在各自的数据集上表现很好，但它们是数据集专用模型——训练时见过什么模式的序列，就只会预测什么模式的序列。真实世界的预测场景常常很窘迫：新城市的客流数据只有三个月、新上线的设备只有几十条记录、某个行业的数据分布和公开基准完全不一样。没有足够历史数据，专用模型根本训练不起来。&lt;/p&gt;
 &lt;p&gt;NLP 和 CV 领域早已给出答案：先在海量数据上做通识教育（预训练），再在具体任务上微调，甚至完全不训练直接使用（零样本）。GPT、CLIP 都是这个范式。Lag-Llama 要做的，就是把这条路搬到时间序列上，具体带来三个价值：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;零样本泛化——面对一个全新的领域/频率，加载预训练权重即可直接预测，不需要任何该领域的训练数据；&lt;/li&gt;
  &lt;li&gt;少样本快速适应——只给 20% 的历史数据做微调，效果就能追上用 100% 数据训练的专用模型（论文实验证实，详见下文）；&lt;/li&gt;
  &lt;li&gt;概率输出开箱即用——预测的不是一个点，而是一整个分布，预测区间直接服务于库存、容量、风险这类需要边界值的决策。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;  &lt;img alt="" height="633" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/lag-llama-2.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;基于 lag 的 token 化：把周期性直接写进输入&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;lag 特征就是回头看的间隔：预测今天 10 点的客流量，最有用的输入不是全部历史，而是昨天 10 点（间隔 24 小时）、上周一 10 点（间隔 168 小时）和上一小时（间隔 1）这几个特定时刻的值。对每个时间步 $t$，模型取过去若干个固定间隔的值组成向量：&lt;/p&gt;
 &lt;p&gt;$$k_t[j] = x_{t – \mathcal{L}[j]}$$&lt;/p&gt;
 &lt;p&gt;其中 $\mathcal{L} = \{\ell_1, \ell_2, \dots\}$ 是一个排好序的 lag 索引集合。举个手算的小例子：小时级序列在 $t$ 时刻，取 $x_{t-1}=102$、$x_{t-2}=98$、$x_{t-24}=110$、$x_{t-168}=95$，那么这一时刻的 lag 特征就是向量 $[102, 98, 110, 95]$。预训练使用的 lag 集合覆盖季度、月、周、日、小时、分钟、秒 7 档频率（由 GluonTS 的 get_lags_for_frequency 按频率自动生成后合并去重），所以同一天/周/年的周期在不同采样频率下都能被表达。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="532" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/lag-llama-3.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;这个设计的核心动机是频率无关性：模型不需要知道数据是小时级还是季度级，只看相对间隔。代价是它需要比上下文窗口更长的历史——实际输入长度 = context_length + 最大 lag，历史太短的序列喂不饱它（这一点我们在局限一节再展开）。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;日期时间特征：让模型读懂时间刻度&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;如果说 lag 特征回答的是多久之前，那么日期时间特征回答的就是这一刻是几点几分、周几、几月。每个 token 额外拼接一组从秒内分钟一直覆盖到年内季度的时间特征。这些特征有个巧妙性质：相邻两个时间步之间，只有一个时间特征发生变化——模型因此能隐式推断出序列的采样频率，从而处理预训练时从未见过的频率组合。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;汇总统计量：告诉模型这条序列的量级&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;如果说 lag 特征给出形状，汇总统计量给出的就是量级。预训练语料里有的序列是几十瓦的传感器读数，有的是几百万的订单量，数值范围差几个数量级。Lag-Llama 在每个窗口上计算中位数 Med 和四分位距 IQR 作为两个时间无关的标量协变量拼进每个 token，相当于告诉模型：这条序列大致多大、波动有多宽。这保证了跨数据集归一化后模型不会晕数字。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;LLaMA 式解码器：RMSNorm + RoPE 的因果 Transformer&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;特征准备好后，交给一个和 LLaMA 同构的解码器：M 层带因果掩码的 Transformer 层，每层使用 RMSNorm 预归一化，并在注意力层的 query/key 上施加旋转位置编码（RoPE）。因果掩码保证只用过去预测未来；RoPE 把相对时间距离编码进注意力，让模型知道两个 token 隔了多远。官方发布的权重规模仅约 245 万参数（论文搜索出的最优配置：8 层、9 个注意力头、每头 16 维、上下文 32 步）——在动辄上百亿参数的大模型时代，它小得几乎可以放进任何环境。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;分布头：输出 Student-t 分布，而不是一个点&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;最后一层把隐藏向量映射为 Student-t 分布的三个参数：自由度 $\nu$、均值 $\mu$、尺度 $\sigma$，训练目标是最小化预测分布的负对数似然（NLL）。预测的不是明天 10 点的客流是 1200，而是它服从一个中心约 1200、带厚尾的分布——这是概率预测的核心思想。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="632" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/student-t.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;为什么选 Student-t 而不是最简单的正态？因为它有厚尾：真实业务序列里常有促销、天气突变、故障这类尖峰，正态分布会把这些当成不可能事件过度惩罚，厚尾则天然宽容。论文刻意选了最简单的参数化分布头（未来可换 normalizing flows、copula 等更复杂的分布），保持模型尽可能简单是作者反复强调的设计原则。&lt;/p&gt;
 &lt;p&gt;推理时采用贪心自回归解码：用预测出的分布采样下一步，把采样值当作已知历史继续往后推，直到预测长度 P 为止；多次采样得到多条未来轨迹，取经验分位数即可得到预测区间。&lt;/p&gt;
 &lt;p&gt;  &lt;strong&gt;鲁棒标准化与数据增强：训练稳定性的两道保险&lt;/strong&gt;&lt;/p&gt;
 &lt;p&gt;标准化公式采用减中位数、除 IQR的鲁棒版本，而不是常见的减均值、除标准差：&lt;/p&gt;
 &lt;p&gt;$$x’_t = \frac{x_t – \operatorname{Med}(x_{1:C})}{\operatorname{IQR}(x_{1:C})}, \qquad \operatorname{IQR}(x_{1:C}) = \operatorname{Med}(\{x_{\lceil C/2 \rceil : C}\}) – \operatorname{Med}(\{x_{1 : \lfloor C/2 \rfloor}\})$$&lt;/p&gt;
 &lt;p&gt;直观地说：一个突然放大 100 倍的尖峰足以把均值、标准差拉爆，但对中位数和四分位距的影响很小，所以异常值不会污染整个窗口的归一化。数据增强方面使用频域增强 Freq-Mask / Freq-Mix（随机遮蔽/混合序列的频率分量），预训练最优增强概率为 0.5，配合按数据集序列总数加权的分层采样，有效缓解跨领域语料上的过拟合。&lt;/p&gt;
 &lt;table width="620"&gt;

  &lt;tr&gt;
   &lt;td width="110"&gt;设计要素&lt;/td&gt;
   &lt;td width="208"&gt;一句话作用&lt;/td&gt;
   &lt;td width="301"&gt;代码/参数入口&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;lag token&lt;/td&gt;
   &lt;td width="208"&gt;注入周期先验，实现频率无关&lt;/td&gt;
   &lt;td width="301"&gt;lags_seq（按频率自动生成）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;日期时间特征&lt;/td&gt;
   &lt;td width="208"&gt;让模型识别采样频率&lt;/td&gt;
   &lt;td width="301"&gt;time_feat=True&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;汇总统计量&lt;/td&gt;
   &lt;td width="208"&gt;传递序列的量级与波动信息&lt;/td&gt;
   &lt;td width="301"&gt;scaling（mean/robust）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;解码器&lt;/td&gt;
   &lt;td width="208"&gt;因果序列建模&lt;/td&gt;
   &lt;td width="301"&gt;n_layer、n_head、n_embd_per_head&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;分布头&lt;/td&gt;
   &lt;td width="208"&gt;概率输出（Student-t）&lt;/td&gt;
   &lt;td width="301"&gt;distr_output=studentT&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="110"&gt;数据增强&lt;/td&gt;
   &lt;td width="208"&gt;缓解过拟合&lt;/td&gt;
   &lt;td width="301"&gt;aug_prob、freq_mask_rate、freq_mix_rate&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;h2&gt;预训练与效果验证&lt;/h2&gt;
 &lt;p&gt;预训练语料是 27 个公开数据集、约 7,965 条单变量序列、约 3.52 亿个窗口 token，覆盖能源、交通、经济、自然、空气质量、云运维六大领域。作者用 catch22 特征（一套 22 维的时序特征描述符）验证了语料覆盖了广泛的时序表现型，这是模型能泛化到未见数据的前提。&lt;/p&gt;
 &lt;table width="632"&gt;

  &lt;tr&gt;
   &lt;td&gt;预训练配置&lt;/td&gt;
   &lt;td width="476"&gt;取值&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据语料&lt;/td&gt;
   &lt;td width="476"&gt;27 个公开数据集，六大领域（能源/交通/经济/自然/空气质量/云运维）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;序列规模&lt;/td&gt;
   &lt;td width="476"&gt;7,965 条单变量序列 ≈ 3.52 亿窗口 token&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;采样策略&lt;/td&gt;
   &lt;td width="476"&gt;分层采样：按数据集总序列数加权&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;数据增强&lt;/td&gt;
   &lt;td width="476"&gt;Freq-Mask / Freq-Mix，增强概率 0.5&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;优化器 / 学习率&lt;/td&gt;
   &lt;td width="476"&gt;Adam / 1e-4&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;Batch size&lt;/td&gt;
   &lt;td width="476"&gt;256&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;每 epoch 窗口数&lt;/td&gt;
   &lt;td width="476"&gt;100（长度 = 最大 lag + 上下文）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;早停&lt;/td&gt;
   &lt;td width="476"&gt;50 epochs（按验证损失）&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;训练硬件&lt;/td&gt;
   &lt;td width="476"&gt;单张 Nvidia Tesla P100（12GB）&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;论文还给出了超参数搜索的最优配置（随机搜索 100 组配置，按预训练验证损失挑选），其中最值得注意的两点：上下文长度只需要 32 步，权重衰减和 Dropout 都设为 0——模型很小，正则化反而多余；以及一个重要的宏观发现：验证损失随预训练数据量按幂律下降，说明时序基础模型同样遵循神经网络的缩放定律，数据与模型继续做大仍有收益。&lt;/p&gt;
 &lt;table width="307"&gt;

  &lt;tr&gt;
   &lt;td width="154"&gt;    &lt;strong&gt;超参数&lt;/strong&gt;&lt;/td&gt;
   &lt;td width="84"&gt;    &lt;strong&gt;搜索范围&lt;/strong&gt;&lt;/td&gt;
   &lt;td width="68"&gt;    &lt;strong&gt;最优值&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;层数 M&lt;/td&gt;
   &lt;td width="84"&gt;1–9&lt;/td&gt;
   &lt;td width="68"&gt;8&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;注意力头数&lt;/td&gt;
   &lt;td width="84"&gt;1–9&lt;/td&gt;
   &lt;td width="68"&gt;9&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;每头嵌入维度&lt;/td&gt;
   &lt;td width="84"&gt;16–512&lt;/td&gt;
   &lt;td width="68"&gt;16&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;上下文长度 C&lt;/td&gt;
   &lt;td width="84"&gt;32–1024&lt;/td&gt;
   &lt;td width="68"&gt;32&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;增强概率&lt;/td&gt;
   &lt;td width="84"&gt;0–1.0&lt;/td&gt;
   &lt;td width="68"&gt;0.5&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;Freq-Mask 比率&lt;/td&gt;
   &lt;td width="84"&gt;0–1.0&lt;/td&gt;
   &lt;td width="68"&gt;0.5&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;Freq-Mix 比率&lt;/td&gt;
   &lt;td width="84"&gt;0–1.0&lt;/td&gt;
   &lt;td width="68"&gt;0.25&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="154"&gt;权重衰减 / Dropout&lt;/td&gt;
   &lt;td width="84"&gt;0–1.0&lt;/td&gt;
   &lt;td width="68"&gt;0 / 0&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;效果验证采用 CRPS（连续排序概率分数，把预测分布与真实值在所有分位数上的差距积分，越低越好）：&lt;/p&gt;
 &lt;p&gt;$$\operatorname{CRPS}(F, x) = \int_{-\infty}^{\infty} \left(F(y) – \mathbb{1}\{y \geq x\}\right)^2 \, dy$$&lt;/p&gt;
 &lt;p&gt;下表是 7 个未见过的数据集上、基于 100 个经验采样样本计算的 CRPS 值，以及模型在 15 个基线中的平均排名：&lt;/p&gt;
 &lt;table width="616"&gt;

  &lt;tr&gt;
   &lt;td width="210"&gt;数据集（领域）&lt;/td&gt;
   &lt;td&gt;零样本 CRPS&lt;/td&gt;
   &lt;td&gt;微调后 CRPS&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;weather（天气）&lt;/td&gt;
   &lt;td&gt;0.164 ± 0.001&lt;/td&gt;
   &lt;td&gt;0.132 ± 0.001&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;ped-counts（行人计数）&lt;/td&gt;
   &lt;td&gt;0.285 ± 0.033&lt;/td&gt;
   &lt;td&gt;0.227 ± 0.010&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;ett-m2（电力变压器温度）&lt;/td&gt;
   &lt;td&gt;0.063 ± 0.002&lt;/td&gt;
   &lt;td&gt;0.017 ± 0.001&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;platform-delay（平台延迟）&lt;/td&gt;
   &lt;td&gt;0.091 ± 0.002&lt;/td&gt;
   &lt;td&gt;0.096 ± 0.002&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;requests（服务器请求量）&lt;/td&gt;
   &lt;td&gt;0.090 ± 0.015&lt;/td&gt;
   &lt;td&gt;0.012 ± 0.002&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;beijing-pm2.5（空气质量）&lt;/td&gt;
   &lt;td&gt;0.130 ± 0.009&lt;/td&gt;
   &lt;td&gt;0.125 ± 0.021&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;exchange（汇率）&lt;/td&gt;
   &lt;td&gt;0.011 ± 0.001&lt;/td&gt;
   &lt;td&gt;0.009 ± 0.000&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td width="210"&gt;平均排名（15 个基线）&lt;/td&gt;
   &lt;td&gt;6.714&lt;/td&gt;
   &lt;td&gt;2.786&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;  &lt;img alt="" height="612" src="https://www.biaodianfu.com/wp-content/uploads/2026/08/crps.png" width="1140"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;两个结论值得记住。第一，零样本已经很能打：不训练任何参数，平均排名 6.714，已经和一大批认真训练过的专用模型处在同一水平（DeepAR 5.714、Informer 6.429），在 weather、platform-delay 等数据集上尤其接近最优。第二，微调才是它的高光：在 ETT-M2、weather、requests 三个数据集上达到 SOTA，平均排名 2.786，领先最强监督基线 TFT（5.000）两个名次以上；更夸张的是少样本场景——只给 20% / 40% / 60% / 80% 的历史数据做微调，平均排名分别达到 1.857 / 1.500 / 1.571 / 1.429，每一档都是全场第一。注意也有反例：platform-delay 微调后 CRPS 不降反升（0.091 → 0.096），说明微调不是万能药，数据量过小时反而可能扰动预训练学到的知识。&lt;/p&gt;
 &lt;h2&gt;上手实践：零样本 → 微调 → 评估&lt;/h2&gt;
 &lt;p&gt;官方实现基于 GluonTS 构建，安装、下载权重、预测的流程非常短。第一步永远是先跑零样本：一个 29MB 的权重文件 + 十行代码，就能拿到带区间的预测。&lt;/p&gt;
 &lt;pre&gt;git clone https://github.com/time-series-foundation-models/lag-llama.git
cd lag-llama
pip install -r requirements.txt

# 下载官方预训练权重（约 29MB，Apache-2.0 许可）
huggingface-cli download time-series-foundation-models/Lag-Llama lag-llama.ckpt --local-dir .
&lt;/pre&gt;
 &lt;pre&gt;import torch
from lag_llama.gluon.estimator import LagLlamaEstimator

# 1) 加载官方预训练权重（lag-llama.ckpt 与脚本同目录）
ckpt = torch.load(lag-llama.ckpt, map_location=cpu)

# 2) 创建估计器：预测未来 24 步，用历史 64 步，采样 100 条轨迹
estimator = LagLlamaEstimator(
    ckpt_path=lag-llama.ckpt,
    prediction_length=24,
    context_length=64,
    nonnegative_pred_samples=True,   # 客流/电量等非负场景建议开启
    num_parallel_samples=100,        # 轨迹数越多，分位数越平滑
)

predictor = estimator.create_predictor(
    estimator.create_transformation(),
    estimator.create_lightning_module(),
)

# 3) 用 GluonTS 内置数据集做零样本预测（示例：澳大利亚用电需求）
from gluonts.dataset.repository.datasets import get_dataset

dataset = get_dataset(australian_electricity_demand, regenerate=False)
test_data = dataset.test

for forecast in predictor.predict(test_data):
    mean    = forecast.mean                  # 预测均值
    p50     = forecast.quantile(0.5)         # 中位数
    lo, hi  = forecast.quantile(0.1), forecast.quantile(0.9)  # 80% 预测区间
零样本确认可用后，再用自己的数据微调。微调的关键是从预训练权重出发、减小模型规模、放宽正则——官方 notebook 的常用配置如下：
from lag_llama.gluon.estimator import LagLlamaEstimator
from gluonts.dataset.common import ListDataset

# 把自己的数据包装成 GluonTS 格式：每条序列一个 {start, target}
train_data = ListDataset(
    [{start: start_ts, target: values}],   # start 为 pd.Timestamp，values 为 np.ndarray
    freq=H,                                   # 数据频率（D/H/T 等）
)

estimator = LagLlamaEstimator(
    ckpt_path=lag-llama.ckpt,   # 从预训练权重出发（迁移学习）
    prediction_length=48,
    context_length=32,
    n_layer=3,                    # 微调时减少层数，防止小数据过拟合
    n_embd_per_head=16,
    n_head=4,
    scaling=mean,
    nonnegative_pred_samples=True,
    aug_prob=0.5,                 # 数据增强开关
    freq_mask_rate=0.1,
    freq_mix_rate=0.05,
    lr=1e-3,                      # 微调用更大学习率
    batch_size=32,
    num_parallel_samples=50,
)

transformation = estimator.create_transformation()
train_dl = estimator.create_train_dataloader(train_data, transformation)
valid_dl = estimator.create_validation_dataloader(valid_data, transformation)

trainer = estimator.create_trainer()
trainer.fit(estimator.create_lightning_module(), train_dl, valid_dl)
最后用 GluonTS 的评估工具量化效果。注意论文里的 CRPS 与分位数加权损失本质等价（CRPS 是对所有分位数的 pinball loss 求积分），因此用 GluonTS 的 Evaluator 即可：
from gluonts.evaluation import make_evaluation_predictions, Evaluator

forecast_it, ts_it = make_evaluation_predictions(
    dataset=test_data,          # GluonTS 格式的测试集
    predictor=predictor,
    num_samples=100,
)

evaluator = Evaluator(quantiles=[0.1, 0.5, 0.9])
agg_metrics, _ = evaluator(ts_it, forecast_it)
print(agg_metrics[mean_wQuantileLoss])   # 分位数加权损失（越低越好）
&lt;/pre&gt;
 &lt;p&gt;实践中还有四个容易踩的坑：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;权重版本——官方在 2024 年中修复了 KV-cache 的因果注意力正确性问题，务必从 Hugging Face 拉取最新lag-llama.ckpt，并在自己的数据上做回归对比；&lt;/li&gt;
  &lt;li&gt;采样数——num_parallel_samples直接决定区间质量：太少（如 10）分位数毛糙，100 是常用值，但推理成本随之线性上升；&lt;/li&gt;
  &lt;li&gt;lag 集合——默认 lag 面向常见周期，如果你的序列有不常见周期（如 30 天账期、学期制），需要自定义lags_seq；&lt;/li&gt;
  &lt;li&gt;单变量限制——每条序列独立预测，多变量问题要么逐通道预测，要么换 Moirai / TimesFM 等多变量方案。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;它最合适的三个场景：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;库存/安全库存——决策需要边界值而非均值，概率输出（分布头 + 采样）直接给出缺货概率，适配原因：区间越准，备货成本越低；&lt;/li&gt;
  &lt;li&gt;运维容量与成本——服务器请求量、云资源用量通常只有少量历史且模式多样，零样本即可上线，微调后逼近专用模型，适配原因：少数据场景正是基础模型的主场；&lt;/li&gt;
  &lt;li&gt;出行/零售的客流需求预测——早晚高峰、周末、节假日的周期性极强，lag 特征天然覆盖日/周/年周期，而大促、突发天气带来的尖峰正好落在 Student-t 厚尾的容忍范围内。&lt;/li&gt;
&lt;/ul&gt;
 &lt;h2&gt;与同期模型的横向对比&lt;/h2&gt;
 &lt;p&gt;Lag-Llama 之后，时序基础模型迅速形成四强并立的格局。它们最大的分歧在于如何把连续序列变成模型能吃的输入：Lag-Llama 用 lag 特征，TimesFM 用 patch 化，Chronos 把数值量化成 token（把时序当语言建模），Moirai 则同时支持多变量与混合频率。&lt;/p&gt;
 &lt;table width="760"&gt;

  &lt;tr&gt;
   &lt;td&gt;维度&lt;/td&gt;
   &lt;td&gt;Lag-Llama&lt;/td&gt;
   &lt;td&gt;TimesFM&lt;/td&gt;
   &lt;td&gt;Chronos&lt;/td&gt;
   &lt;td&gt;Moirai&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;提出时间&lt;/td&gt;
   &lt;td&gt;2023-10（论文）/ 2024-02（权重）&lt;/td&gt;
   &lt;td&gt;2024&lt;/td&gt;
   &lt;td&gt;2024-03&lt;/td&gt;
   &lt;td&gt;2024-05&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;输入表示&lt;/td&gt;
   &lt;td&gt;lag 特征 + 日历特征&lt;/td&gt;
   &lt;td&gt;patch 化&lt;/td&gt;
   &lt;td&gt;量化 token&lt;/td&gt;
   &lt;td&gt;patch + 任意变量&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;概率输出&lt;/td&gt;
   &lt;td&gt;Student-t 分布头&lt;/td&gt;
   &lt;td&gt;分位数头&lt;/td&gt;
   &lt;td&gt;token 分类分布采样&lt;/td&gt;
   &lt;td&gt;分布输出&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;变量类型&lt;/td&gt;
   &lt;td&gt;单变量&lt;/td&gt;
   &lt;td&gt;单变量（可加协变量）&lt;/td&gt;
   &lt;td&gt;单变量&lt;/td&gt;
   &lt;td&gt;任意变量&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;参数量&lt;/td&gt;
   &lt;td&gt;≈2.45M&lt;/td&gt;
   &lt;td&gt;≈200M&lt;/td&gt;
   &lt;td&gt;8M–710M（5 档）&lt;/td&gt;
   &lt;td&gt;small / base / large&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;最擅长&lt;/td&gt;
   &lt;td&gt;轻量、易微调、原生概率&lt;/td&gt;
   &lt;td&gt;长上下文、零样本&lt;/td&gt;
   &lt;td&gt;把时序当语言建模&lt;/td&gt;
   &lt;td&gt;多变量、MoE&lt;/td&gt;
&lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;许可证&lt;/td&gt;
   &lt;td&gt;Apache-2.0&lt;/td&gt;
   &lt;td&gt;Apache-2.0&lt;/td&gt;
   &lt;td&gt;Apache-2.0&lt;/td&gt;
   &lt;td&gt;Apache-2.0&lt;/td&gt;
&lt;/tr&gt;

&lt;/table&gt;
 &lt;p&gt;选型建议可以一句话概括：要原生概率输出、要在小数据上快速微调、预算有限——选 Lag-Llama；要超长上下文零样本——选 TimesFM；要处理多变量联动——选 Moirai；想体验把时序当语言的新范式——看 Chronos。&lt;/p&gt;
 &lt;h2&gt;局限与破局&lt;/h2&gt;
 &lt;p&gt;基础模型不是银弹，Lag-Llama 的边界非常清晰，四条局限都对应着明确的破局路径：&lt;/p&gt;
 &lt;ul&gt;
  &lt;li&gt;只做单变量——不建模多个变量之间的相关性，多变量系统只能逐通道预测、丢失联动信息。破局：Moirai 等任意变量模型直接建模多变量；Lag-Llama 自身也有引入 LSTM 增强长期依赖的变体研究；&lt;/li&gt;
  &lt;li&gt;lag 特征要求最小上下文——实际输入必须 ≥ 最大 lag，短序列（如只有一周数据的冷启动场景）喂不饱。破局：Chronos 的量化 token 和 TimesFM 的 patch 化对上下文要求更宽松；极短序列也可回到 ETS / ARIMA 等传统方法；&lt;/li&gt;
  &lt;li&gt;零样本只是能打，不是最强——真正的 SOTA 要靠微调，且个别数据集微调后反而变差（如 platform-delay）。破局：预算内优先微调，用官方 notebook 配置起步；用验证集做早停，防止小数据过拟合破坏预训练知识；&lt;/li&gt;
  &lt;li&gt;生态更新节奏慢——2024 年中才修复 KV-cache 正确性问题，checkpoint 演进不如大厂模型频繁。破局：锁定版本 + 建立回归测试；需要长期维护的生产系统可评估 TimesFM 等更新更活跃的选项。&lt;/li&gt;
&lt;/ul&gt;
 &lt;p&gt;Lag-Llama 用lag 特征 + LLaMA 解码器 + Student-t 分布头这个朴素组合，证明了时序基础模型这条路走得通，并由此开启了 2024 年时序基础模型的全面爆发——今天的 TimesFM、Chronos、Moirai，都是在回答它提出的同一个问题：如何让一个模型学会预测所有时间序列。&lt;/p&gt;
 &lt;div&gt;

  &lt;strong&gt;相关文章:&lt;/strong&gt;  &lt;ol&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/moirai/" rel="bookmark" title="&amp;#26102;&amp;#38388;&amp;#24207;&amp;#21015;&amp;#22522;&amp;#30784;&amp;#27169;&amp;#22411;Moirai 2.0"&gt;时间序列基础模型Moirai 2.0&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/hive-sql-guide/" rel="bookmark" title="Hive SQL&amp;#31995;&amp;#32479;&amp;#21270;&amp;#23398;&amp;#20064;"&gt;Hive SQL系统化学习&lt;/a&gt;&lt;/li&gt;
   &lt;li&gt;    &lt;a href="https://www.biaodianfu.com/llama-cpp/" rel="bookmark" title="&amp;#26412;&amp;#22320;&amp;#21270;&amp;#37096;&amp;#32626;&amp;#22823;&amp;#27169;&amp;#22411;&amp;#24037;&amp;#20855; llama.cpp"&gt;本地化部署大模型工具 llama.cpp&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>器→工具 工具软件 数据 术→技巧 LLM</category>
      <guid isPermaLink="true">https://itindex.net/detail/63275-llm-%E5%BC%80%E6%BA%90-%E9%A2%84%E6%B5%8B</guid>
      <pubDate>Mon, 31 Aug 2026 22:09:03 CST</pubDate>
    </item>
    <item>
      <title>高性价比大模型GLM 5.3 Flash 和 Qwen3.8-Flash 同一天发布</title>
      <link>https://itindex.net/detail/63274-%E6%80%A7%E4%BB%B7%E6%AF%94-%E6%A8%A1%E5%9E%8B-glm</link>
      <description>&lt;div&gt;  &lt;div&gt;   &lt;ul&gt;    &lt;li&gt;     &lt;strong&gt;Ox Alpha 正式模型名称为 GLM 5.3 Flash Vison&lt;/strong&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/div&gt;  &lt;div&gt;
划重点：&lt;/div&gt;  &lt;div&gt;
- 国产芯片雄起，这几天全球每天超过 100T 的供应量，全部由国产芯片提供算力支持，低价管饱，改变行业格局&lt;/div&gt;  &lt;div&gt;
- Artificial Analysis Intelligence Index 57 分，和 Opus 4.8 打平&lt;/div&gt;  &lt;div&gt;
- 模型定价，现在限时半价，只要 Opus 4.8 的 1/40，新的性价比斩杀线已经形成&lt;/div&gt;  &lt;div&gt;
- 超低成本的全新架构，参数 320B，激活参数 18 B，价格做到了上一代 GLM 5.3 的十分之一&lt;/div&gt;  &lt;div&gt;
- 真正原生多模态模型，不止是能看见，还能感知环境，通过视觉迭代改进，可以自己在 Blender 画出一个厨房。。。&lt;/div&gt;  &lt;div&gt;

GLM 5.3 Flash 目前已经在    &lt;a href="https://t.co/HALAezrowT" rel="noopener noreferrer nofollow" target="_blank"&gt;http://z.ai&lt;/a&gt; 上线&lt;/div&gt;  &lt;div&gt;
   &lt;a href="https://t.co/KCaX1FvRqz" rel="noopener noreferrer nofollow" target="_blank"&gt;http://colaos.ai&lt;/a&gt; 已经在上周五接入限免测试版，明天正式上线正式版&lt;/div&gt;  &lt;div&gt;

价格和接入：&lt;/div&gt;  &lt;div&gt;
   &lt;a href="https://t.co/E877shCMjR" rel="noopener noreferrer nofollow" target="_blank"&gt;https://docs.bigmodel.cn/cn/guide/models/vlm/glm-5.3-flash…&lt;/a&gt;&lt;/div&gt;  &lt;div&gt;

开源：&lt;/div&gt;  &lt;div&gt;
   &lt;a href="https://t.co/ssWg9rastc" rel="noopener noreferrer nofollow" target="_blank"&gt;https://huggingface.co/zai-org/GLM-5.3-Flash…&lt;/a&gt;&lt;/div&gt;  &lt;div&gt;   &lt;br /&gt;&lt;/div&gt;  &lt;div&gt;   &lt;br /&gt;&lt;/div&gt;  &lt;div&gt;   &lt;br /&gt;&lt;/div&gt;  &lt;div&gt;   &lt;ul&gt;    &lt;li&gt;     &lt;strong&gt;Qwen3.8-Flash&lt;/strong&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/div&gt;  &lt;div&gt;一个多模态 MoE 模型，同时也是 Qwen4 架构的早期预览版，现已开源权重！&lt;/div&gt;  &lt;div&gt;

生产版本 Qwen3.8-Flash 即将通过 QwenCloud API 提供，价格仅为 $ 0.16/1M 输入令牌 和 $ 0.47/1M 输出令牌。&lt;/div&gt;  &lt;div&gt;

125B 参数 + 51B N-gram 嵌入，每个令牌仅激活 6B。无与伦比的性价比。&lt;/div&gt;  &lt;div&gt;

新特性：   &lt;img alt="&amp;#55358;&amp;#56691;" src="https://abs.twimg.com/emoji/v2/svg/1f973.svg" title="Partying face"&gt;&lt;/img&gt;&lt;/div&gt;  &lt;div&gt;
- 下一代架构：GDN + QSA 混合注意力、门控残差、N-gram 嵌入 &amp;amp; Muon 优化器，作为 Qwen4 中使用的架构的前身。&lt;/div&gt;  &lt;div&gt;
- 训练和推理成本大幅降低：训练成本仅为 Qwen3.7-Plus 的 1/9，同时在各方面超越它，尤其在编码和办公任务中表现出色。&lt;/div&gt;  &lt;div&gt;
- 强劲性能：在 DeepSWE 1.1 上得分 58.7，在 SWE-bench Pro 上得分 62.5，在 CoWorkBench 上得分 73.9，在 AndroidWorld 上得分 84.5，在 MathVision（带 CI）上得分 95.7。&lt;/div&gt;  &lt;div&gt;
- 262K 原生上下文，可通过 YaRN 扩展至 1M。&lt;/div&gt;  &lt;div&gt;

Qwen还发布了 Qwen3.8-Flash-Next 的权重

- 博客：   &lt;a href="https://t.co/M5hYypFLgJ" rel="noopener noreferrer nofollow" target="_blank"&gt;https://qwen.ai/blog?id=qwen3.8-flash-next…&lt;/a&gt;&lt;/div&gt;  &lt;div&gt;
- 技术报告：   &lt;a href="https://t.co/IF0gObIkQO" rel="noopener noreferrer nofollow" target="_blank"&gt;https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf…&lt;/a&gt;&lt;/div&gt;  &lt;div&gt;
- Hugging Face：   &lt;a href="https://t.co/6ow8QVAABt" rel="noopener noreferrer nofollow" target="_blank"&gt;https://huggingface.co/Qwen/Qwen3.8-Flash-Next?spm=a2ty_o06.30285417.0.0.1d73c921FsyOPe&amp;amp;file=Qwen3.8-Flash-Next…&lt;/a&gt;&lt;/div&gt;  &lt;div&gt;
   &lt;br /&gt;&lt;/div&gt;&lt;/div&gt; &lt;div&gt;  &lt;div&gt;   &lt;div&gt;    &lt;div&gt;     &lt;div&gt;      &lt;div&gt;       &lt;div&gt;        &lt;a href="https://x.com/oran_ge/status/2092620850252177504/photo/1"&gt;         &lt;br /&gt;         &lt;div&gt;&lt;/div&gt;&lt;/a&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63274-%E6%80%A7%E4%BB%B7%E6%AF%94-%E6%A8%A1%E5%9E%8B-glm</guid>
      <pubDate>Thu, 27 Aug 2026 11:24:21 CST</pubDate>
    </item>
    <item>
      <title>刚刚，新款 Mac mini 发布！价格大涨2500元，苹果第一次为 AI 造电脑</title>
      <link>https://itindex.net/detail/63273-mac-mini-%E4%BB%B7%E6%A0%BC</link>
      <description>&lt;p&gt; &lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1307" src="https://s3.ifanr.com/wp-content/uploads/2026/08/111-2.jpg" width="1960"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;2026 年最好的 AI PC，刚刚迎来了一大波更新。&lt;/p&gt;
 &lt;p&gt;苹果正式发布全新一代 Mac mini 和 Mac Studio。新一代桌面 Mac 带来了 M6、M5 Pro、M5 Max 与 M5 Ultra 四款芯片选择，也进一步拉高了苹果在本地 AI 计算领域的性能上限。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1307" src="https://s3.ifanr.com/wp-content/uploads/2026/08/111-2.jpg" width="1960"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;其中，Mac mini 提供 M6 和 M5 Pro 两种版本，起售价分别为 6999 元和 12999 元。针对教育用户，苹果还提供对应优惠价格，M6 版 Mac mini 起售价为 6199 元，M5 Pro 版为 12199 元。&lt;/p&gt;
 &lt;p&gt;Mac Studio 则提供 M5 Max 与 M5 Ultra 两种版本，起售价分别为 19999 元和 46999 元，将模型容量、内存带宽和多机扩展推向更高层级。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1448" src="https://s3.ifanr.com/wp-content/uploads/2026/08/1-11.png" width="1086"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;作为对比，M4 Mac mini 国行起售价在今年 6 月已经由 4499 元涨至 5999 元，M4 Pro 版目前为 12499 元起。新款 M6 和 M5 Pro 版又分别上涨 1000 元和 500 元。&lt;/p&gt;
 &lt;p&gt;四款全新机型均于 8 月 27 日开启预购，9 月 22 日起发售。其中，配备 512GB 统一内存的 Mac Studio 将于 10 月下旬上市。&lt;/p&gt;
 &lt;h3&gt;苹果最小的 Mac，正在成为 AI 时代的新入口&lt;/h3&gt;
 &lt;p&gt;过去，Mac mini 在苹果产品线里的定位非常明确。它是一台价格相对友好的 Mac，让用户用较低成本进入 macOS 生态。&lt;/p&gt;
 &lt;p&gt;但新款 Mac mini 的任务伴随着年初 OpenClaw 的爆火，它除了承担办公、影音和日常创作，也逐渐成为越来越多用户运行本地 AI 模型和 AI Agent 的设备选择。&lt;/p&gt;
 &lt;p&gt;具体而言，新款 Mac mini 的核心变化，来自全新的 M6 芯片。&lt;/p&gt;
 &lt;p&gt;作为苹果首款采用 2 纳米制程的 M 系列芯片，M6 配备 12 核 CPU、12 核 GPU，以及两组 16 核神经网络引擎。其中，CPU 包含 2 个超级核心、4 个性能核心和 6 个能效核心，核心总数比 M5 增加 2 个。&lt;/p&gt;
 &lt;p&gt;苹果官方给出的数据显示，M6 的多线程 CPU 性能最高达到 M5 的 1.2 倍、M1 的 2.4 倍。放到 Mac mini 中，与上一代 M4 版本相比，其 CPU 性能最高提升 40%，图形性能最高达到 2 倍，AI 性能最高达到 4 倍。&lt;/p&gt;
 &lt;p&gt;M6 的 12 核 GPU 首次在 Mac mini 的每个 GPU 核心中加入神经网络加速器。其 AI 峰值计算能力较 M5 提升接近 30%，较 M1 提升超过 8 倍，主要用于加快大模型的提示词处理。&lt;/p&gt;
 &lt;p&gt;两组 16 核神经网络引擎的峰值计算能力最高达到前代的 2 倍。系统框架可以同时调用两组引擎，提高设备端模型的执行效率。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1491" src="https://s3.ifanr.com/wp-content/uploads/2026/08/2-8.png" width="1055"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;M6 版从 16GB 统一内存起步，最高可选 32GB，内存带宽达到每秒 170GB，较 M5 提高 10%，是 M1 的 2.5 倍。&lt;/p&gt;
 &lt;p&gt;在具体应用中，这款机型使用 LM Studio 处理大模型提示词的速度，最高达到 M1 版 Mac mini 的 13.5 倍、M4 版的 4.8 倍。&lt;/p&gt;
 &lt;p&gt;其 Microsoft Excel 表格计算速度最高达到 M1 版的 2.3 倍、M4 版的 1.5 倍。在支持光线追踪的《赛博朋克 2077：终极版》中，游戏性能最高达到 M4 版的 2 倍。&lt;/p&gt;
 &lt;p&gt;由此来看，M6 版 Mac mini 已经不只是入门办公电脑。&lt;/p&gt;
 &lt;p&gt;它同时面向普通用户、学生、开发者、AI 爱好者和企业用户，可以处理智能编码、图像生成、模型推理和轻量级 Agent 任务。&lt;/p&gt;
 &lt;p&gt;如果你有更复杂的工作负载，苹果还准备了 M5 Pro 版本。&lt;/p&gt;
 &lt;p&gt;M5 Pro 最高配备 18 核 CPU 和 20 核 GPU，每个 GPU 核心同样包含神经网络加速器。统一内存最高达到 64GB，带宽达到每秒 307GB，可以容纳体量更大的模型、复杂三维场景、ProRes RAW 视频文件和科研数据集。&lt;/p&gt;
 &lt;p&gt;在 LM Studio 中，M5 Pro 版 Mac mini 的大模型提示词处理速度最高达到 M2 Pro 版的 8.5 倍、M4 Pro 版的 4 倍。其 Blender 光线追踪渲染性能最高达到 M2 Pro 版的 4.5 倍，Affinity 图像处理性能最高达到 2.1 倍。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1402" src="https://s3.ifanr.com/wp-content/uploads/2026/08/3-5.png" width="1122"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;苹果官方介绍还将其描述为适合 AI Agent 或创意工作流的安静、高效常驻设备。这一定位说明，Mac mini 的使用场景正在从用户主动操作，延伸到后台持续运行。&lt;/p&gt;
 &lt;p&gt;连接能力也围绕这一方向升级。&lt;/p&gt;
 &lt;p&gt;两种版本均支持 Wi-Fi 7、蓝牙 6 和 2.5Gb 以太网，并可选配 10Gb 以太网。机身正面配备两个支持 USB 3 的 USB-C 接口和一个高阻抗耳机插孔。&lt;/p&gt;
 &lt;p&gt;M6 版背面配备三个 Thunderbolt 4 接口，M5 Pro 版则配备三个 Thunderbolt 5 接口，另外还有 HDMI 和以太网接口。&lt;/p&gt;
 &lt;p&gt;借助 Thunderbolt 5，多台 M5 Pro 版 Mac mini 可以组成集群，共同承载超出单机能力的模型。新增的 USB-C 同步锁相功能，则可以让显示设备与 iPhone 17 Pro 等拍摄设备实现精确同步。&lt;/p&gt;
 &lt;p&gt;至此，Mac mini 获得了一个新身份。它既是一台个人桌面电脑，也可以成为连接数据、应用和 Agent 的小型 AI 节点。体型虽小，野心拉满。&lt;/p&gt;
 &lt;h3&gt;从单机到集群，Mac Studio 组队越级&lt;/h3&gt;
 &lt;p&gt;如果说 Mac mini 提供了个人 AI 计算的入口，那么 Mac Studio 承担的就是更高强度的模型推理、训练与创意工作。&lt;/p&gt;
 &lt;p&gt;新款 Mac Studio 提供 M5 Max 与 M5 Ultra 两种版本。&lt;/p&gt;
 &lt;p&gt;M5 Max 配备 18 核 CPU，其中包括 6 个超级核心和 12 个性能核心。GPU 最高拥有 40 个核心，每个核心都内置神经网络加速器。统一内存最高达到 128GB，带宽达到每秒 614GB。&lt;/p&gt;
 &lt;p&gt;与 M4 Max 相比，M5 Max 在 LM Studio 中处理大模型提示词的速度最高达到 3.9 倍，文生图性能最高达到 3.5 倍，DaVinci Resolve Studio 的 Magic Mask 性能最高达到 3 倍。与更早的 M1 Max 相比，其大模型提示词处理速度最高达到 10.7 倍。&lt;/p&gt;
 &lt;p&gt;到了 M5 Ultra，苹果干脆把「力大砖飞」写进了芯片结构。&lt;/p&gt;
 &lt;p&gt;它通过新一代 UltraFusion 技术，将两颗采用双芯片晶粒设计的 M5 Max 连接起来，形成苹果首个四晶粒 M 系列系统级芯片。&lt;/p&gt;
 &lt;p&gt;UltraFusion 的晶粒间带宽超过每秒 4.4TB，连接密度提升超过 6 倍，使四个晶粒可以像一个统一处理器一样运行。&lt;/p&gt;
 &lt;p&gt;M5 Ultra 最高配备 36 核 CPU 和 80 核 GPU，并首次在 Ultra 芯片的每个 GPU 核心中加入神经网络加速器。其 AI 峰值性能最高达到 M3 Ultra 的 4.3 倍、M1 Ultra 的 9.8 倍。按照芯片层面的峰值 GPU AI 算力计算，其性能最高达到 M3 Ultra 的 4.5 倍。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1586" src="https://s3.ifanr.com/wp-content/uploads/2026/08/4-5.png" width="992"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;搭载 M5 Ultra 的 Mac Studio 最高支持 512GB 统一内存，带宽达到每秒 1.2TB。这样的容量可以在设备内保存大型数据集，并承载数百亿乃至数千亿参数的模型。&lt;/p&gt;
 &lt;p&gt;在 LM Studio 中，该机型的大模型提示词处理速度最高达到 M1 Ultra 版的 9.8 倍、M3 Ultra 版的 4 倍，文生图性能最高达到 8.2 倍和 4.3 倍；Foundry Nuke 的 CopyCat 训练速度最高达到 M1 Ultra 版的 15.4 倍。&lt;/p&gt;
 &lt;p&gt;M5 Ultra 还配备更强的媒体引擎，可以同时播放最多 33 路每秒 30 帧的 8K ProRes 422 视频。新款 Mac Studio 的存储速度最高提升至 2 倍，采用基于 PCIe 6 的新一代固态存储架构。&lt;/p&gt;
 &lt;p&gt;Mac Studio 首次支持 Wi-Fi 7 和 Bluetooth 6，同时加入 Thunderbolt 5 连接能力，可以连接高速外置存储、PCIe 扩展机箱和其他专业设备。整机最多支持 8 台显示器，或者 4 台以 5K 分辨率、120Hz 刷新率运行的 Studio Display XDR。&lt;/p&gt;
 &lt;p&gt;相比单机性能，多机集群更能体现这次升级对产品逻辑的改变。&lt;/p&gt;
 &lt;p&gt;新款 Mac Studio 支持通过 Thunderbolt 5 和远程直接内存访问技术连接。多台设备可以组成低延迟网络，共同执行分布式 AI 推理。根据苹果公布的数据，4 机集群的推理性能最高可以达到单机的 3 倍。&lt;/p&gt;
 &lt;p&gt;值得一提的是，今年 6 月，苹果在 WWDC26 期间的一场特别活动中，展示了这种集群的实际运行方式。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="713" src="https://s3.ifanr.com/wp-content/uploads/2026/08/5-5.png" width="1280"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;现场的 4 台 Mac Studio 通过 Thunderbolt 5 连接，并利用 RDMA over Thunderbolt 建立低延迟通信网络。&lt;/p&gt;
 &lt;p&gt;借助 LM Studio 当时即将推出、基于 MLX Distributed 的功能，LM Studio 员工加载并运行了 Kimi K2.6 这一万亿参数规模的开放权重模型。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="713" src="https://s3.ifanr.com/wp-content/uploads/2026/08/6-7.png" width="1280"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;其背后的关键是分布式计算。&lt;/p&gt;
 &lt;p&gt;对于一万亿参数模型，如果采用 FP16 精度，仅模型权重理论上就需要约 2TB 内存，尚未计入缓存和其他运行开销。单台个人电脑很难提供如此大的可用空间，多机协同则可以同时扩展计算能力、内存容量与带宽。&lt;/p&gt;
 &lt;p&gt;不过，桌面集群的实际性能仍然受到模型精度、内存配置、通信速度和模型架构影响。&lt;/p&gt;
 &lt;h3&gt;下一代个人电脑的用户，未必是人类&lt;/h3&gt;
 &lt;p&gt;今年 2 月 6 日，可能是人类最后一次在 AI 服务中消耗了比 AI Agent 更多的 token。&lt;/p&gt;
 &lt;p&gt;OpenRouter 数据显示，此后 Agent 的 token 使用量超过人类，并在约半年内增长至原来的 14 倍，人类当然并没有减少使用 AI，只是越来越多 token 的触发指令被交给了 Agent 代劳。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="1254" src="https://s3.ifanr.com/wp-content/uploads/2026/08/7-5.png" width="1254"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;因 OpenClaw 爆火而受到关注的 Mac mini，也展示了一种此前并不突出的硬件需求。&lt;/p&gt;
 &lt;p&gt;以前，人敲打键盘，电脑执行，然后等待下一次指令。Agent 的到来改变了这种供需关系。&lt;/p&gt;
 &lt;p&gt;一台面向 Agent 的设备，需要保留用户的文件、历史记录、应用权限、模型状态和任务上下文。它不仅要完成一次推理，还要知道此前做过什么、当前进行到哪一步，以及接下来可以调用哪些资源。&lt;/p&gt;
 &lt;p&gt;从这个角度看，以 Mac mini 为代表的 Mac 电脑受到关注，不只是因为它体积小、功耗较低或者性能足够强。更深层的原因是，它具备成为这种数字环境的基本条件。&lt;/p&gt;
 &lt;p&gt;它可以接入用户的本地文件和系统应用，也能将敏感数据保留在设备内。&lt;/p&gt;
 &lt;p&gt;不过，外形和能效只能解释它为什么适合常驻桌面，真正决定模型运行能力的，是 Apple Silicon 的统一内存架构。&lt;/p&gt;
 &lt;p&gt;传统电脑通常由 CPU 使用系统内存，独立 GPU 使用显存。&lt;/p&gt;
 &lt;p&gt;模型权重和计算数据需要通过总线在两类内存之间传输。独立显卡虽然拥有很强的峰值算力，但消费级产品的显存通常只有几十 GB，模型规模很容易受到容量限制。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="827" src="https://s3.ifanr.com/wp-content/uploads/2026/08/8-4.png" width="1080"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;Apple Silicon 让 CPU、GPU、神经网络引擎和其他模块共享同一个物理内存池。模型权重、输入数据和中间结果只需要保存一份，减少了重复存储和数据搬运。&lt;/p&gt;
 &lt;p&gt;M6 最高支持 32GB 统一内存，带宽为每秒 170GB。M5 Pro 将两项指标提高至 64GB 和每秒 307GB。M5 Max 进一步达到 128GB 和每秒 614GB。M5 Ultra 则提供最高 512GB 容量与每秒 1.2TB 带宽。&lt;/p&gt;
 &lt;p&gt;统一内存负责容纳模型，芯片提供计算能力，Core AI、Core ML、Metal 和 MLX 等框架负责调度硬件，操作系统与本地应用则向 Agent 提供数据和工具。&lt;/p&gt;
 &lt;p&gt;这些能力结合在一起，构成了 Mac 在 Agent 时代的优势。它提供的不只是一次模型推理，而是一套覆盖上下文、权限、状态和执行过程的本地 AI 环境。&lt;/p&gt;
 &lt;p&gt;  &lt;img alt="" height="853" src="https://s3.ifanr.com/wp-content/uploads/2026/08/9-1.jpg" width="1280"&gt;&lt;/img&gt;&lt;/p&gt;
 &lt;p&gt;电脑角色的变化，也开始影响 Mac 的产品分层。&lt;/p&gt;
 &lt;p&gt;过去，苹果主要按照性能需求、工作负载、预算和便携性划分产品线。在这套体系中，mini、Pro、Max 和 Ultra，不同程度地代表着用户对性能的需求，以及所从事工作的专业程度。&lt;/p&gt;
 &lt;p&gt;当电脑的操作者开始包含 AI，新一代桌面 Mac 因而显露出另一套分档逻辑。&lt;/p&gt;
 &lt;p&gt;M6 版 Mac mini 面向日常 AI 工具、智能编码和轻量级 Agent 任务。M5 Pro 版可以承载更复杂的模型与内容处理。M5 Max 版 Mac Studio 进入专业 AI 开发、图像视频生成和大型数据集处理。M5 Ultra 版则指向前沿模型推理、本地训练与多机协同。&lt;/p&gt;
 &lt;p&gt;某种程度上，mini、Pro、Max 和 Ultra 正在逐渐代表一台机器可以承载的 AI 工作负载等级。&lt;/p&gt;
 &lt;p&gt;这不是苹果第一次强化 Mac 的相关性能，却可能是 AI 第一次如此完整地影响 Mac 的整机定位、层级划分和扩展方式。&lt;/p&gt;
 &lt;p&gt;而当 Agent 消耗的 token 超过人类，电脑行业也必须重新面对一个基础问题：一台电脑究竟在为谁工作？新一代 Mac 给出的答案是，它仍然属于人类，但使用它的已经不只有人类。&lt;/p&gt;
 &lt;p&gt;#欢迎关注爱范儿官方微信公众号：爱范儿（微信号：ifanr），更多精彩内容第一时间为您奉上。&lt;/p&gt;&lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category>公司 AI Mac mini 苹果</category>
      <guid isPermaLink="true">https://itindex.net/detail/63273-mac-mini-%E4%BB%B7%E6%A0%BC</guid>
      <pubDate>Tue, 25 Aug 2026 21:10:31 CST</pubDate>
    </item>
    <item>
      <title>为什么你的本地LLM用起来感觉比实际更“笨”</title>
      <link>https://itindex.net/detail/63272-llm-%E6%84%9F%E8%A7%89</link>
      <description>&lt;h2&gt;快速引言&lt;/h2&gt; &lt;p&gt;我们都有过这样的经历：在论坛、聊天、Reddit、Discord、YouTube等地方，听到有人惊呼“哦！某某模型简直太棒了！”，然后下载了它（或者更可能是它的某个量化版本），一试之后却觉得“呃…这太烂了！”&lt;/p&gt; &lt;p&gt;这篇文章将是一系列相当技术性的实验，旨在展示推理过程中，因具体实现方式（implementation-specific）而产生的差异所造成的影响。我会用“参考实现”这个词来指代那些发布模型并提供官方托管、并公布原始基准测试结果的实验室。他们的硬件和软件都会与你的大不相同。而本文中的对比，也不会是在Ollama中跑几个测试提示词、比较2.58-bit的GGUF模型那么简单。&lt;/p&gt; &lt;p&gt;为了让大家更容易理解，我会故意略过一些新兴的研究领域和大量研究论文。请不要对我的过度简化吹毛求疵，否则我可能会让你去读那又长又烦人、还带数学公式的版本。&lt;/p&gt; &lt;h2&gt;你的本地实现很糟糕。但没关系，因为其他人的也一样。&lt;/h2&gt; &lt;p&gt;如今，每一个运行LLM的硬件和软件实例都略有不同，在某些情况下甚至可能大相径庭。一般的家庭实验室用户可能混合使用着不同世代的GPU，这些芯片上的指令集各不相同。这些指令集在执行数学运算来计算你的下一个token时，方式也会因人而异，即使大家运行的是完全相同的权重。&lt;/p&gt; &lt;p&gt;这就引出了第一个问题：你的特定配置到底有多“糟糕”？ 事实证明，有很多不同的方法可以衡量这一点。&lt;/p&gt; &lt;p&gt;实际的方法很直接：运行标准的基准测试，并且要多样化。比如terminal bench、hle、SWEthis、HELLAthat、MMLU-whatever……随便选。关键是要确保这些测试能代表你的实际工作负载和用例。 不要只是把温度调到0，粘贴3个测试提示词，然后就断定模型好或不好。零样本（Zero-shot）测试并不能很好地模拟大多数智能体（agentic）任务。你需要长上下文的工具调用和特定领域的知识评估，才能在与他人运行相同权重并复现相同基准测试时，找出你配置的弱点所在。&lt;/p&gt; &lt;p&gt;但我将从纯粹的数学角度开始讨论，因为正如@wendell所说：&lt;/p&gt; &lt;h2&gt;数学就是数学！&lt;/h2&gt; &lt;p&gt;“Logits（逻辑值）”是模型对每个可能的下一个token给出的评分。这些评分会被归一化为概率，然后通过配置好的采样器（sampler），最后由解码器（detokenizer）转换回文本，从而在解码过程中生成 THE→NE→XT→TOK→EN。&lt;/p&gt; &lt;p&gt;关于采样器设置的一个旁注：Hugging Face模型卡通常会明确指定你应该使用的采样器设置（以及聊天模板），例如 temp 1.0, top-p 0.95 等，不同模型有所差异，请确保你使用了正确的设置。顺便说一句，温度设得太低就是导致你的Qwen模型陷入循环、无法跳出其“THINK”输出的原因。不用谢，很高兴能帮你解决这个问题。&lt;/p&gt; &lt;p&gt;当下一个token的概率发生足够大的变化时，THE→NE→XT 就可能变成 THE→NE→W→DAY……虽然这些微小的变化可能没问题，但很可能是你潜意识里总觉得“哪里不对劲”的源头。&lt;/p&gt; &lt;p&gt;你们有些人可能听过KLD（KL散度，Kullback-Leibler Divergence）这个术语。别担心，我不会让你做数学题，也不会用一堆小数表格轰炸你的大脑。简单来说就是：将输出的logits转换为概率分布，然后衡量该分布相对于选定基线移动了多少。较低的KLD值意味着更接近基线，但并不自动代表模型“更聪明”。 KLD也是有方向的，所以两个分布的顺序很重要。&lt;/p&gt; &lt;p&gt;提醒一句： 不要被某个量化模型HF模型卡上低得不可思议的KLD声明所迷惑。除非作者披露了参考检查点、完整的运行时环境、评估文本、校准数据、上下文长度、采样位置、KL方向、任何词汇表截断以及测量结果的聚合方式，否则这个数字是无法解读的。方法论和数字本身同样重要，而很多人在这方面会出错。&lt;/p&gt; &lt;h2&gt;vllm到底在做什么？&lt;/h2&gt; &lt;p&gt;现在，我们需要简要探讨一下你的推理引擎上那庞大的软件栈到底在做什么，以便理解这些差异的一些来源。&lt;/p&gt; &lt;p&gt;在这个被极度简化的流程图中，每一步都包含可以根据你的具体硬件/软件配置、模型、量化方式、张量形状等进行配置或更改的组件。&lt;/p&gt; &lt;p&gt;我抓取的vLLM每日构建版（nightly）容器镜像包含了734个软件包。这意味着有734个代码库，每个都有其自身的错误和未文档化的特性。你的特定实现穿过这座“代码山”的路径将是独一无二的。&lt;/p&gt; &lt;h1&gt;测试1：注意力后端的精度基准测试&lt;/h1&gt; &lt;p&gt;让我们从推理流程图中的一部分开始。在预填充（prefill，即提示词处理）阶段，你的推理引擎会从几个注意力后端中选择一个。这会影响预填充的速度和精度，同时不同的GPU家族/流处理器（SM）计算能力版本需要不同的CUDA内核。我们来测试并比较它们。&lt;/p&gt; &lt;p&gt;（非常抱歉，但我觉得有必要这么做…）&lt;/p&gt; &lt;p&gt;我从Qwen3.6-27B的官方BF16检查点开始，在RTX PRO 6000 Blackwell GPU上运行，张量并行度（tensor parallelism）设为1。KV缓存使用BF16格式，没有采用权重、激活或KV缓存的量化。软件环境是固定版本的vLLM每日构建版。我使用了eager执行模式，禁用了CUDA图、前缀缓存和MTP，并采用了2k-token的分块预填充（chunked prefill）。&lt;/p&gt; &lt;p&gt;Qwen3.6-27B是密集（dense）模型，而非MoE（混合专家）模型，但它仍是一个混合架构：64层以“三个门控DeltaNet/线性注意力层”加“一个全注意力层”的模式重复。只有那16个全注意力层在此实验中使用可选注意力后端；门控DeltaNet路径保持不变。&lt;/p&gt; &lt;p&gt;这里重放的工作负载是“Prompt 2”，一个约10万token的上下文，来自真实的Turnstone实验室工作流，包含多次工具调用和实际工作产物。选择它是因为它模拟了本地智能体实际的工作方式，而不是那种合成的“大海捞针”测试。而且更重要的是，这个工作负载没有出现在当今任何基准测试或训练数据集中。没有人能针对这个进行过基准优化，或校准他们的量化方案来适应它。&lt;/p&gt; &lt;p&gt;在这个工作负载下，vLLM中有三个可用的全注意力后端可供选择：FlashAttention 2、Flash Inference 和 Triton Attention。这是执行之间唯一更改的设置，其余硬件和软件栈保持稳定。&lt;/p&gt; &lt;p&gt;我还进行了相同后端、跨GPU的可重复性控制测试。在这张图表中，我每32个提示词token捕获一次BF16格式的完整词汇表logits。像KLD这样的分布比较，是之后在FP64精度下从这些存储的logits中计算得出的。&lt;/p&gt; &lt;p&gt;“Top-1一致性”是指具有最高logit的token（即贪婪解码的argmax）是否相同。所有三个后端都针对相同的强制token历史进行了评估。因此，“top-1翻转”意味着在该位置，某个后端会选择不同的贪婪解码下一个token。我们并没有让这个选择改变后续的历史。这保证了数学比较的可控性，但它无法展示一个不受约束的生成过程会如何分支，或者一次工具调用是否会最终失败……这部分将在测试2中揭晓 ;D&lt;/p&gt; &lt;p&gt;下图显示了导致token翻转的采样logits百分比：&lt;/p&gt; &lt;p&gt;在最初的几千个token中，无论使用哪种后端，模型对下一个token的预测都是一致的。但在提示词的后半部分，后端开始出现分歧。这里选择Triton作为基准，以简化后续的量化对比。&lt;/p&gt; &lt;p&gt;每个8k-token窗口包含250个采样位置（每32个token探测一次）。百分比指的是在这些探测点中，其他后端得分最高的token与Triton不同的比例。&lt;/p&gt; &lt;p&gt;通过使用相同的注意力后端多次运行相同测试，来排除随机噪声的影响。结果显示，每次运行的logits在每个隐藏状态上都逐位（bit for bit）相同。这意味着观察到的差异完全来自于在预填充期间，在Triton/FA2/FI内部发生的矩阵乘法和加法运算。&lt;/p&gt; &lt;p&gt;分歧呈集群式出现，并且随提示词内容变化，而非随上下文长度平滑增加。这并不能证明存在某个使模型“崩溃”的通用长度阈值……但我们很快就会讲到……&lt;/p&gt; &lt;p&gt;现在我们有了一个关于特定提示词工作负载的基线比较，让我们深入探讨……&lt;/p&gt; &lt;h1&gt;测试2：KV缓存量化，或为什么你的LLM在超过4万token后智商陡降&lt;/h1&gt; &lt;p&gt;重复相同的方法，我们以上述运行Triton的BF16权重和BF16 KV缓存基线为基准，进行了下一个实验：保持权重和激活不变，仅量化KV缓存，会发生什么？&lt;/p&gt; &lt;p&gt;答案就是：分歧。这引出了我们今晚的第一个“大型翻车现场”：一个完全可重现的工具调用错误。&lt;/p&gt; &lt;p&gt;在工具调用期间，足够多的top-token被翻转了。我们让它们继续执行，结果发现：BF16一切正常，int8 KV缓存最终设法恢复了，而 int4 KV缓存则完全无法恢复！&lt;/p&gt; &lt;h1&gt;测试3：权重，别告诉我！&lt;/h1&gt; &lt;p&gt;这次我们保持所有KV缓存为完整的BF16大小。但我们引入了一些新的选手来进行比较：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;BF16参考：Qwen/Qwen3.6-27B 官方&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;官方FP8：Qwen/Qwen3.6-27B-FP8 官方&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;INT8 W8A16：TheHouseOfTheDude 社区量化&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;NVIDIA NVFP4：nvidia/Qwen3.6-27B-NVFP4 官方&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;AWQ W4A16：cyankiwi 社区量化&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;这4种量化方案代表了权重和激活的广泛图景。对我们“数学爱好者”来说，一个值得注意的信息是，计算每种量化logits时所实际运行的CUDA内核 / GEMM（通用矩阵乘法） / MMA（矩阵乘加）指令都是不同的：&lt;/p&gt; &lt;h3&gt;Qwen3.6-27B (参考)&lt;/h3&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;权重/激活：BF16权重，BF16激活&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;线性层/GEMM：未量化线性方法 → torch.nn.functional.linear。每个CUDA块由其相关的形状/几何结构选择。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV缓存：BF16（强制）&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;特性：参考检查点。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;h3&gt;Qwen3.6-27B-FP8&lt;/h3&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;权重/激活：128×128块中的E4M3 FP8权重；在转换后的线性层内进行动态FP8激活量化；    &lt;code&gt;lm_head&lt;/code&gt;等排除模块保持BF16。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;线性层/GEMM：Fp8LinearMethod → CutlassFp8BlockScaledMMKernel&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV缓存：BF16（强制）&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;特性：由于vLLM将其E8M0缩放格式标记为此架构（SM120）的精度降级，DeepGemm被自动禁用；转而选择了CUTLASS。发布文件中未识别出校准数据集。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;h3&gt;Qwen3.6-27B-INT8&lt;/h3&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;权重/激活：静态、对称、通道级别的INT8线性权重；BF16激活 (W8A16)。GDN/linear_attn投影和    &lt;code&gt;lm_head&lt;/code&gt;被排除在量化之外。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;线性层/GEMM：CompressedTensorsWNA16 → MarlinLinearKernel&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV缓存：BF16（强制）&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;特性：一次性量化（one-shot），明确没有使用校准数据集。考虑到W8A16加上未量化的GDN投影，其异常良好的保真度就不那么神秘了。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;h3&gt;Qwen3.6-27B-NVFP4&lt;/h3&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;权重/激活：混合检查点——覆盖64个全注意力投影和144个GDN投影的208个静态FP8 W8A8目标；覆盖192个MLP投影加上    &lt;code&gt;lm_head&lt;/code&gt;的193个NVFP4 W4A16目标，组大小16。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;线性层/GEMM：&lt;/p&gt;   &lt;ul&gt;    &lt;li&gt;     &lt;p&gt;FP8目标：ModelOptFp8LinearMethod → FlashInferFP8ScaledMMLinearKernel&lt;/p&gt;&lt;/li&gt;    &lt;li&gt;     &lt;p&gt;NVFP4目标：NVFP4 GEMM → MarlinNvFp4LinearKernel&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV缓存：BF16（强制）&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;特性：在我们的上游每日构建版运行中，并非原生FP4算术。vLLM判定该GPU路径缺乏原生FP4支持，并明确通过Marlin选择了纯权重的FP4压缩。检查点中嵌入的FP8 KV方案在本次对比中被BF16 KV覆盖。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;h3&gt;Qwen3.6-27B-AWQ-BF16-INT4&lt;/h3&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;权重/激活：静态非对称INT4权重，组大小32，MSE观察器；BF16激活 (W4A16)。GDN/linear_attn投影和    &lt;code&gt;lm_head&lt;/code&gt;被排除。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;线性层/GEMM：CompressedTensorsWNA16 → MarlinLinearKernel&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;KV缓存：BF16（强制）&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;特性：AWQ校准数据集被披露为“STEM和智能体（STEM and Agentic）”。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;本次运行的其它值得注意信息：&lt;/p&gt; &lt;ul&gt;  &lt;li&gt;   &lt;p&gt;所有模型的完整softmax/GQA注意力均为     &lt;code&gt;AttentionBackendEnum.TRITON_ATTN&lt;/code&gt;；JIT监视器观察到     &lt;code&gt;kernel_unified_attention&lt;/code&gt;。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;GDN预填充：Triton/FLA GDN预填充内核，请求为     &lt;code&gt;triton&lt;/code&gt;，    &lt;code&gt;head_k_dim=128&lt;/code&gt;。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;执行期间，循环路径还JIT编译了     &lt;code&gt;_causal_conv1d_update_kernel&lt;/code&gt;，    &lt;code&gt;fused_recurrent_gated_delta_rule_packed_decode_kernel&lt;/code&gt; 和     &lt;code&gt;reduce_segments&lt;/code&gt;。&lt;/p&gt;&lt;/li&gt;  &lt;li&gt;   &lt;p&gt;张量并行度1，eager模式，无CUDA图，无MTP/推测解码，仅语言执行。&lt;/p&gt;&lt;/li&gt;&lt;/ul&gt; &lt;p&gt;下一个token翻转的结果相当可预测。TheDude的INT8 (W8A16) 方案表现最佳，击败了第一方的FP8 (W8A8) 和NVIDIA（FP4名不副实）的发布版。事实上，在5个选项中，NVIDIA的发布版排名垫底，在上下文达到88k时，其token翻转率达到了约50%。&lt;/p&gt; &lt;p&gt;在工具调用方面，NVFP4和AWQ W4A16都未能正确关闭它们的工具调用，并且搞乱了Cisco的命令行语法（正确命令是&amp;apos;show arp&amp;apos;，而它们执行了&amp;apos;show run&amp;apos;），而FP8和INT8都能成功完成正确的调用。&lt;/p&gt; &lt;p&gt;在未来的实验中，我将尝试探索对相同权重使用不同融合GEMM的影响，这是另一个有趣的差异来源，有时你不得不在精度和速度之间做出取舍。&lt;/p&gt; &lt;div&gt;  &lt;br /&gt;&lt;/div&gt;
     
    &lt;div&gt; &lt;a href="https://itindex.net/"  title="IT 资讯"&gt;&lt;img src="https://itindex.net/images/iconWarning.gif" title="IT 资讯" border="0"/&gt; &lt;/a&gt;</description>
      <category />
      <guid isPermaLink="true">https://itindex.net/detail/63272-llm-%E6%84%9F%E8%A7%89</guid>
      <pubDate>Mon, 24 Aug 2026 11:24:11 CST</pubDate>
    </item>
  </channel>
</rss>


