<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="zh-CN">
<title>墨痕</title>
<subtitle>关于线性注意力、语言模型与工程实践的个人笔记。</subtitle>
<link href="https://yaunt.github.io/blog"/>
<link href="https://yaunt.github.io/blog/atom.xml" rel="self"/>
<id>https://yaunt.github.io/blog</id>
<updated>2026-10-10T16:45:01Z</updated>
<generator>Inkstone Static Blog Engine</generator>
<entry><title>Exact Linear Attention</title><link href="https://yaunt.github.io/blog/blogs/exact-linear-attention/app.html"/><id>https://yaunt.github.io/blog/blogs/exact-linear-attention/app.html</id><updated>2026-06-04T00:00:00Z</updated><published>2026-06-04T00:00:00Z</published><author><name>Yauntyours</name></author><category term="线性注意力"/><category term="注意力机制"/><category term="MoE"/><category term="记忆模块"/><category term="YOLO"/><category term="视觉模型"/><category term="核函数"/><summary type="text">线性注意力长期被视为把 Transformer 扩展到长序列的关键方向。本文给出一种&quot;精确&quot;线性注意力：利用核函数的精确可分解性，把 $O(L^2)$ 注意力改写为 $O(L)$ 的线性形式而不损失精度，并进一步讨论 FFN 可解释性、Hyper-Link 残差替换与&quot;质记忆&quot;模块，最后把方法推广到视觉模型。</summary><content type="html"><![CDATA[<p align="center">
<a href="https://arxiv.org/abs/2605.18848" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/badge/arXiv-2605.18848-b31b1b.svg?style=flat-square&logo=arxiv&logoColor=white" alt="arXiv"></a>&nbsp;
<a href="https://github.com/yauntyour/Exact-Linear-Attention/stargazers" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/github/stars/yauntyour/Exact-Linear-Attention?style=flat-square&logo=github" alt="GitHub Stars"></a>&nbsp;
<a href="https://github.com/yauntyour/Exact-Linear-Attention/network/members" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/github/forks/yauntyour/Exact-Linear-Attention?style=flat-square&logo=github" alt="GitHub Forks"></a>&nbsp;
<a href="https://github.com/yauntyour/Exact-Linear-Attention/issues" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/github/issues/yauntyour/Exact-Linear-Attention?style=flat-square&logo=github" alt="GitHub Issues"></a>&nbsp;
<a href="https://github.com/yauntyour/Exact-Linear-Attention/pulls" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/github/issues-pr/yauntyour/Exact-Linear-Attention?style=flat-square&logo=github" alt="GitHub Pull Requests"></a>&nbsp;
<a href="https://github.com/yauntyour/Exact-Linear-Attention/commits/main" target="_blank" rel="noopener noreferrer"><img loading="lazy" decoding="async" src="https://img.shields.io/github/last-commit/yauntyour/Exact-Linear-Attention?style=flat-square&logo=github" alt="Last Commit"></a>
</p>
<h2 id="table-of-contents">Table of Contents</h2>
<ul>
<li><a href="#1-introduction">1. Introduction</a></li>
<li><a href="#2-regarding-which-kernel-function-is-suitable-for-this-task">2. Regarding Which Kernel Function is Suitable for This Task</a></li>
<li><a href="#3-exact-linear-attention-formulation">3. Exact Linear Attention Formulation</a></li>
<li><a href="#4-how-to-construct-your-own-attention-kernel">4. How to Construct Your Own Attention Kernel</a></li>
<li><a href="#5-challenges-in-engineering-construction">5. Challenges in Engineering Construction</a>
<ul>
<li><a href="#51-the-dilemma-of-ffn">5.1 The Dilemma of FFN</a></li>
<li><a href="#52-hyperlink-based-residual-replacement-for-degradation-eradication">5.2 Hyperlink-Based Residual Replacement for Degradation Eradication</a></li>
<li><a href="#53-how-memory-works">5.3 How Memory Works</a></li>
</ul>
</li>
<li><a href="#6-experiments">6. Experiments</a>
<ul>
<li><a href="#61-no-memory-training-comparison">6.1 No Memory Training Comparison</a></li>
<li><a href="#62-memory-additional">6.2 Memory Additional</a></li>
<li><a href="#63-long-range-training">6.3 Long-range Training</a></li>
</ul>
</li>
<li><a href="#7-extend-to-vision-models">7. Extend to Vision Models</a></li>
<li><a href="#8-exploratory-works">8. Exploratory Works</a></li>
<li><a href="#acknowledgments">Acknowledgments</a></li>
<li><a href="#references">References</a></li>
</ul>
<hr />
<h2 id="1-introduction">1. Introduction</h2>
<p>Linear attention has long been considered a promising direction for scaling Transformers to long sequences.</p>
<p>Inspired by the studies of linear attention <a href="#ref-katharopoulos2020">(Katharopoulos et al., 2020)</a> <a href="#ref-schlag2021">(Schlag et al., 2021)</a> from Katharopoulos' group and Jürgen Schmidhuber's team, this paper formally presents an exact linear attention mechanism. It exploits the exact decomposition property of kernels, departing from the traditional approximate softmax paradigm.</p>
<p>Meanwhile, after systematically analyzing the limitations of linear attention, the MiniMax Team identified two key issues: gradient explosion induced by denominator-caused unbounded gradients, and token attention dilution in linear attention <a href="#ref-qin2022">(Qin et al., 2022)</a>. Targeting these defects, we propose a linearized full attention computation approach, which bounds gradients by imposing kernel constraints and exploits the inherent favorable properties of kernels. This work offers a new perspective to resolve the inherent limitations of linear attention.</p>
<p>By breaking the <span class="math math--inline" data-tex="O(L^2)">O(L^2)</span> computational bottleneck of conventional attention, linear attention enables dramatic improvements in computational efficiency without the need for sampling. When the context length increases by one order of magnitude, the performance gap between linear and conventional attention widens correspondingly. Consequently, linearized attention will continue to be an essential research pursuit for Transformer-style attention networks <a href="#ref-vaswani2017">(Vaswani et al., 2017)</a> over the long term.</p>
<p>Based on empirical results from teams including Kimi Linear (The Dark Side of the Moon) <a href="#ref-kimilinear2025">(Kimi Team, 2025)</a> and MiniMax <a href="#ref-minimax2025lightning">(MiniMax Team, 2025)</a>, this theoretical complexity gap manifests as a substantial performance disparity in practical applications:</p>
<ul>
<li><strong>Inference Speed (Throughput):</strong> For short texts (e.g., 4k context length), the speed difference between the two approaches is marginal, and linear attention may even be slightly slower due to operator optimization limitations. In ultra-long text scenarios (e.g., 1M tokens), the decoding speed of linear attention exceeds that of full global attention by more than six times. In a 1M context test, Kimi Linear achieves a Time Per Output Token (TPOT) of merely 1.84 ms, while traditional architectures such as MLA reach as high as 11.48 ms.</li>
<li><strong>Memory Usage (KV Cache):</strong> Linear attention reduces the KV Cache memory footprint by 75%–90%.</li>
</ul>
<p>This indicates that with identical GPU resources, linear attention models can process much longer documents or serve more concurrent users with a larger batch size. For instance, MiniMax's model leverages linear attention to support a context window of up to 4 million tokens, a scale that is prohibitively costly for conventional full-attention architectures.</p>
<p><img src="assets/Introduce.webp" alt="Introduction Figure" width="1536" height="1024" srcset="assets/Introduce-820.webp 820w, assets/Introduce.webp 1536w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<hr />
<h2 id="2-regarding-which-kernel-function-is-suitable-for-this-task">2. Regarding Which Kernel Function is Suitable for This Task</h2>
<p>To address the inherent limitations of linear attention, we summarize the desirable properties that an ideal kernel function should possess.</p>
<ul>
<li><strong>Exactly decomposable</strong> — A kernel function shall have the inherent property to admit expansions of adequate accuracy.</li>
<li><strong>Sufficiently discriminative</strong> — The output curve is expected to possess prominent discriminability along with a smooth and broad value range, thereby mitigating the issues of gradient vanishing and gradient explosion.</li>
<li><strong>Non‑negative</strong> — Guaranteeing that all attention matrix weights are non‑negative.</li>
<li><strong>Geometrically interpretable</strong> — The kernel ought to possess a clear geometric interpretation in the embedding space, so as to enhance the internal interpretability of the model.</li>
</ul>
<p>The above points cover the most fundamental characteristics required for attention computation. Exact decomposability guarantees the precision of attention calculation; sufficient discriminability ensures that the resulting attention matrix will not be diluted after normalization. Meanwhile, non-negativity and geometric interpretability serve as the basic prerequisites for attention models.</p>
<p>Through systematic summarization and induction, we find that kernel functions satisfying the requirements can be roughly divided into the following categories: Polynomial type, exponential type, non-negative periodic function type, and absolute value function type.</p>
<p>Given the demand for favorable discriminability and nonlinear characteristics, the Hadamard Exp Kernel stands out as an optimal choice for the following reasons:</p>
<ol>
<li>It is exactly decomposable and maintains a relatively low computational complexity.</li>
<li>The exponential transformation provides strong nonlinearity and enhances feature selectivity.</li>
<li>It is fully smooth and differentiable everywhere.</li>
<li>It guarantees non-negativity of attention weights.</li>
</ol>
<p>These desirable properties allow us to easily associate the essence of attention — query-based cosine-like similarity — and thereby recognize that the following kernel functions possess distinct geometric interpretability.</p>
<ul>
<li>
<p><strong>Summation Squared Euclidean Distance Kernel:</strong> Its attention mechanism emphasizes keys that are aligned in the same direction as the query. For example, it retrieves supporting evidence in question answering, namely contextual content consistent with the query direction.</p>
<div class="math math--block" data-tex="\|A_i + B_j\|^2 = \|A\|^2 + \|B\|^2 + 2A_i\cdot B_j">\|A_i + B_j\|^2 = \|A\|^2 + \|B\|^2 + 2A_i\cdot B_j</div>
</li>
<li>
<p><strong>Subtraction Squared Euclidean Distance Kernel:</strong> Its attention mechanism emphasizes keys that are opposite or antagonistic to the query. Typical applications include contrastive learning and the detection of contradictory paragraphs.</p>
<div class="math math--block" data-tex="\|A_i - B_j\|^2 = \|A\|^2 + \|B\|^2 - 2A_i\cdot B_j">\|A_i - B_j\|^2 = \|A\|^2 + \|B\|^2 - 2A_i\cdot B_j</div>
</li>
<li>
<p><strong>Hadamard Exp Kernel:</strong> Its attention mechanism emphasizes feature co-activation patterns through exponential amplification. Examples are given as follows:</p>
<ul>
<li>Multimodal scenarios: Different modalities activate the same semantic features.</li>
<li>Feature selection: Attention acts as a feature gate, amplifying strongly activated features.</li>
<li>Noise robustness: The exponential operation suppresses low-intensity noise while enhancing salient signals.</li>
</ul>
<div class="math math--block" data-tex="k(A_i,B_j) = \exp(A_i) \ast \exp(B_j) = \sum_{d=1}^{D} \exp(A_{id})\exp(B_{jd})">k(A_i,B_j) = \exp(A_i) \ast \exp(B_j) = \sum_{d=1}^{D} \exp(A_{id})\exp(B_{jd})</div>
</li>
</ul>
<p>Most other kernel variants are derivatives of these three categories, which will not be elaborated further here. Of particular note compared with canonical attention is the <strong>Hadamard Exp Kernel</strong>. The element-wise exponential product characterizes the co-occurrence intensity across feature dimensions, where the exponential transformation naturally amplifies strongly activated feature pairs while suppressing noise. Compared with the cosine-similarity-based paradigm of conventional attention, it can capture feature co-activation patterns in a more discriminative manner, presenting clear advantages in scenarios requiring fine-grained feature interaction. The same applies to the summation and subtraction Euclidean distance kernels, which capture magnitude-directional relationships. In contrast, the Hadamard Exp Kernel is especially suitable for multimodal scenarios where semantic feature co-activation across different modalities is critical.</p>
<hr />
<h2 id="3-exact-linear-attention-formulation">3. Exact Linear Attention Formulation</h2>
<p>We now derive the attention formula following the standard attention paradigm.</p>
<p>Let <span class="math math--inline" data-tex="A \in \mathbb{R}^{B \times L \times D}">A \in \mathbb{R}^{B \times L \times D}</span> and <span class="math math--inline" data-tex="B \in \mathbb{R}^{B \times L \times D}">B \in \mathbb{R}^{B \times L \times D}</span> be the query and key representations, and let <span class="math math--inline" data-tex="V \in \mathbb{R}^{B \times L \times d_v}">V \in \mathbb{R}^{B \times L \times d_v}</span> be the value matrix. To simplify the formulation in our discussion, we denote <span class="math math--inline" data-tex="A \in \mathbb{R}^{B \times L \times D}">A \in \mathbb{R}^{B \times L \times D}</span> as <span class="math math--inline" data-tex="A_i">A_i</span>, where the index <span class="math math--inline" data-tex="i">i</span> corresponds to the dimension <span class="math math--inline" data-tex="L">L</span>. Accordingly, <span class="math math--inline" data-tex="A">A</span> can be regarded as an embedding matrix of size <span class="math math--inline" data-tex="B \times L">B \times L</span>. The same definition applies to <span class="math math--inline" data-tex="B_j">B_j</span> and <span class="math math--inline" data-tex="V_j">V_j</span>.</p>
<p>By virtue of Mercer's theorem <a href="#ref-mercer1909">(Mercer, 1909)</a>, any positive definite kernel admits a decomposition as an inner product within a feature space.</p>
<div class="math math--block" data-tex="k(A_i, B_j) = \sum_{m=1}^{\infty}\lambda_m \phi_m(A_i)\psi_m(B_j)^\top">k(A_i, B_j) = \sum_{m=1}^{\infty}\lambda_m \phi_m(A_i)\psi_m(B_j)^\top</div>
<p>However, we do not require such a fully positive definite decomposition property here. It is sufficient for the kernel function operation on <span class="math math--inline" data-tex="A_i">A_i</span> and <span class="math math--inline" data-tex="B_j">B_j</span> to be decomposed into the product of two sub-kernels.</p>
<div class="math math--block" data-tex="k(A_i, B_j) = \phi(A_i)\psi(B_j)^\top">k(A_i, B_j) = \phi(A_i)\psi(B_j)^\top</div>
<p>The decomposition of the aforementioned kernel functions can be illustrated as follows:</p>
<ul>
<li>
<p><strong>Summation Squared Euclidean Distance Kernel</strong>:</p>
<div class="math math--block" data-tex="k(A_i,B_j) = \|A_i + B_j\|^2 = \|A_i\|^2 + \|B_j\|^2 + 2A_i\cdot B_j">k(A_i,B_j) = \|A_i + B_j\|^2 = \|A_i\|^2 + \|B_j\|^2 + 2A_i\cdot B_j</div>
</li>
<li>
<p><strong>Subtraction Squared Euclidean Distance Kernel</strong>:</p>
<div class="math math--block" data-tex="k(A_i,B_j) = \|A_i - B_j\|^2 = \|A_i\|^2 + \|B_j\|^2 - 2A_i\cdot B_j">k(A_i,B_j) = \|A_i - B_j\|^2 = \|A_i\|^2 + \|B_j\|^2 - 2A_i\cdot B_j</div>
</li>
<li>
<p><strong>Hadamard Exp Kernel</strong>:</p>
<div class="math math--block" data-tex="k(A_i,B_j) = \exp(A_i) \ast \exp(B_j) = \sum_{d=1}^{D} \exp(A_{id})\exp(B_{jd})">k(A_i,B_j) = \exp(A_i) \ast \exp(B_j) = \sum_{d=1}^{D} \exp(A_{id})\exp(B_{jd})</div>
</li>
</ul>
<p>For further illustration, the exact decomposition itself actually imposes little requirement on symmetry.</p>
<div class="math math--block" data-tex="k(A_i, B_j) = A_{i1} B_{j2} + 2 A_{i2} B_{j1}">k(A_i, B_j) = A_{i1} B_{j2} + 2 A_{i2} B_{j1}</div>
<div class="math math--block" data-tex="\phi(A_i) = \begin{pmatrix} A_{i1} \\ 2A_{i2} \end{pmatrix} \in \mathbb{R}^{2},\quad \psi(B_j) = \begin{pmatrix} B_{j2} \\ B_{j1} \end{pmatrix} \in \mathbb{R}^{2}">\phi(A_i) = \begin{pmatrix} A_{i1} \\ 2A_{i2} \end{pmatrix} \in \mathbb{R}^{2},\quad \psi(B_j) = \begin{pmatrix} B_{j2} \\ B_{j1} \end{pmatrix} \in \mathbb{R}^{2}</div>
<p>Clearly, <span class="math math--inline" data-tex="k(A_i, B_j) \neq k(B_j, A_i)">k(A_i, B_j) \neq k(B_j, A_i)</span> in general, showing that the decomposition <span class="math math--inline" data-tex="k(A,B)=\langle\phi(A),\psi(B)\rangle">k(A,B)=\langle\phi(A),\psi(B)\rangle</span> does not require the kernel to be symmetric.</p>
<p>However, it can be clearly recognized that the kernel operation <span class="math math--inline" data-tex="k(A_i, B_j)">k(A_i, B_j)</span> itself is decomposed into two components <span class="math math--inline" data-tex="\phi(A_i)">\phi(A_i)</span> and <span class="math math--inline" data-tex="\psi(B_j)">\psi(B_j)</span>. This allows us to swap their order to implement attention computation with linear complexity <strong>without any loss of precision</strong>.</p>
<div class="math math--block" data-tex="k(A_i, B_j)V_j = \phi(A_i) \psi(B_j)^\top V_j = \phi(A_i)\,[\psi(B_j)^\top V_j]">k(A_i, B_j)V_j = \phi(A_i) \psi(B_j)^\top V_j = \phi(A_i)\,[\psi(B_j)^\top V_j]</div>
<p>On this basis, we perform row normalization on the kernel function <span class="math math--inline" data-tex="k(A_i, B_j)">k(A_i, B_j)</span> to make it conform to a certain probability distribution. In this way, we achieve normalization of the attention distribution while eliminating the need for the softmax operation. Combined with a special mathematical summation operation, the entire process maintains linear computational complexity.</p>
<div class="math math--block" data-tex="\frac{\sum_{j=1}^{L} k(A_i, B_j)V_j}{\sum_{j=1}^{L} k(A_i, B_j)} = \frac{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top V_j}{\sum_{j=1}^{L} \phi(A_i)\psi(B_j)^\top} = \frac{\phi(A_i)\bigl[\sum_{j=1}^{L}\psi(B_j)^\top V_j\bigr]}{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top}">\frac{\sum_{j=1}^{L} k(A_i, B_j)V_j}{\sum_{j=1}^{L} k(A_i, B_j)} = \frac{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top V_j}{\sum_{j=1}^{L} \phi(A_i)\psi(B_j)^\top} = \frac{\phi(A_i)\bigl[\sum_{j=1}^{L}\psi(B_j)^\top V_j\bigr]}{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top}</div>
<p>In practical implementation, we often adopt the following optimization strategies for both bidirectional and causal versions of linear attention:</p>
<ul>
<li>
<p><strong>Bidirectional Attention</strong></p>
<div class="math math--block" data-tex="C = \sum_{j=1}^{L} \psi(B_j), \qquad S = \sum_{j=1}^{L} \psi(B_j) V_j^\top">C = \sum_{j=1}^{L} \psi(B_j), \qquad S = \sum_{j=1}^{L} \psi(B_j) V_j^\top</div>
<div class="math math--block" data-tex="Y_i = \frac{\phi(A_i)^\top S}{\phi(A_i)^\top C}">Y_i = \frac{\phi(A_i)^\top S}{\phi(A_i)^\top C}</div>
</li>
<li>
<p><strong>Causal (Auto-Regressive) Attention</strong></p>
<div class="math math--block" data-tex="C_i = \sum_{j=1}^{i} \psi(B_j), \qquad S_i = \sum_{j=1}^{i} \psi(B_j) V_j^\top">C_i = \sum_{j=1}^{i} \psi(B_j), \qquad S_i = \sum_{j=1}^{i} \psi(B_j) V_j^\top</div>
<div class="math math--block" data-tex="Y_i = \frac{\phi(A_i)^\top S_i}{\phi(A_i)^\top C_i}">Y_i = \frac{\phi(A_i)^\top S_i}{\phi(A_i)^\top C_i}</div>
</li>
</ul>
<p>By swapping the order of summation, the bidirectional version requires only a single accumulation over the sequence, and the causal version uses a prefix sum (cumulative sum). In both cases the entire attention output is computed in <span class="math math--inline" data-tex="O(L)">O(L)</span> time without ever materializing the <span class="math math--inline" data-tex="L \times L">L \times L</span> attention matrix.</p>
<p>Because the kernel is exactly decomposable into finite‑dimensional feature maps, the result is mathematically identical to the full quadratic form—this is an <strong>exact</strong>, rather than approximate, linear attention mechanism.</p>
<hr />
<h2 id="4-how-to-construct-your-own-attention-kernel">4. How to Construct Your Own Attention Kernel</h2>
<p>At this point, I believe you are already eager to get started. Nevertheless, there is no need to rush. Based on the four criteria we have proposed, you can freely design a kernel function tailored to your specific task. It is only necessary to satisfy these four requirements to construct a brand-new attention kernel, which can further achieve the time complexity of <span class="math math--inline" data-tex="O(L^2)">O(L^2)</span> via linearized computation.</p>
<p>For example, if we aim to restore standard attention with the highest possible precision, we can regard its scaled dot-product followed by softmax as first computing the dot product, then performing exponential transformation, and finally conducting normalization. This formulation is equivalent to the exponential dot-product kernel. However, evaluating the exponential dot-product kernel inevitably requires the Taylor expansion formula. Therefore, we instead seek to construct a specialized kernel function that satisfies the exact decomposition condition, while preserving the inherent characteristics of standard attention in capturing both vector magnitude and directional information.</p>
<p>It looks like this:</p>
<div class="math math--block" data-tex="k(A_i, B_j) = (\vec{A}_i \cdot \vec{B}_j + 1) \cdot (\|A_i\|^2 + 1) \cdot (\|B_j\|^2 + 1)">k(A_i, B_j) = (\vec{A}_i \cdot \vec{B}_j + 1) \cdot (\|A_i\|^2 + 1) \cdot (\|B_j\|^2 + 1)</div>
<div class="math math--block" data-tex="\phi(A_i) = (\|A_i\|^2 + 1)\begin{pmatrix} \vec{A}_i \\ 1 \end{pmatrix}, \quad \psi(B_j) = (\|B_j\|^2 + 1)\begin{pmatrix} \vec{B}_j \\ 1 \end{pmatrix}">\phi(A_i) = (\|A_i\|^2 + 1)\begin{pmatrix} \vec{A}_i \\ 1 \end{pmatrix}, \quad \psi(B_j) = (\|B_j\|^2 + 1)\begin{pmatrix} \vec{B}_j \\ 1 \end{pmatrix}</div>
<div class="math math--block" data-tex="\phi(A_i)^\top \psi(B_j) = (\|A_i\|^2 + 1)(\|B_j\|^2 + 1)(\vec{A}_i \cdot \vec{B}_j + 1)">\phi(A_i)^\top \psi(B_j) = (\|A_i\|^2 + 1)(\|B_j\|^2 + 1)(\vec{A}_i \cdot \vec{B}_j + 1)</div>
<p>There is no need to marvel at its complexity. In fact, this kernel function can be viewed as two components. The term <span class="math math--inline" data-tex="\vec{A}_i \cdot \vec{B}_j + 1">\vec{A}_i \cdot \vec{B}_j + 1</span> captures the attention to directional information, while the remaining part <span class="math math--inline" data-tex="(\|A_i\|^2 + 1) \cdot (\|B_j\|^2 + 1)">(\|A_i\|^2 + 1) \cdot (\|B_j\|^2 + 1)</span> accounts for the attention to magnitude information. Following this paradigm, one can theoretically construct arbitrary types of attention kernel functions.</p>
<hr />
<h2 id="5-challenges-in-engineering-construction">5. Challenges in Engineering Construction</h2>
<p>In fact, modern machine learning toolkits such as PyTorch already enable us to rapidly construct ideal model architectures. Nevertheless, the advancement of AI toward AGI is still hindered by issues including communication overhead, memory consumption, energy cost, and even human-related factors. Meanwhile, we notice that all these challenges can be resolved through productivity liberation driven by technological progress. In the following, we focus on several key aspects from an engineering perspective.</p>
<h3 id="5-1-the-dilemma-of-ffn">5.1 The Dilemma of FFN</h3>
<p>At present, the limitation of FFN lies in its poor interpretability, namely the so-called &quot;<strong>black-box</strong>&quot; problem <a href="#ref-jain2019attention">(Jain &amp; Wallace, 2019)</a>. To trace this mapping process, existing studies mostly summarize the statistical patterns and characteristics of pre-trained models. However, such approaches are largely ineffective for MoE models. Due to the sparse activation property of MoE, there exists an inherent gradient inconsistency gap between the router and the expert groups. Under training with hard-constrained load balancing, each expert is trained almost independently. It is impossible to clearly interpret why a specific set of tokens activates a certain expert. Simply attributing this phenomenon to the stronger processing capability of an expert for a particular type of tokens is rather far-fetched and hardly universally accepted. Moreover, the routing allocation mechanism of MoE imposes considerable communication overhead. We do not deny that MoE itself is of remarkable significance for expanding the knowledge capacity of neural networks. Therefore, we aim to develop a method that can perform nonlinear transformation on the semantic embedding space vectors of post-attention outputs without relying on explicit routing dispatching.</p>
<p>Clearly, an attention query mechanism is indispensable. Traditional full attention is avoided due to its prohibitive computational overhead, yet the paradigm has now shifted. Full attention computation with linear complexity offers a viable solution to bridge the semantic gap caused by sparse activation. We assign each expert network a fixed, learnable &quot;<strong>label vector</strong>&quot;, which are aggregated into a unified key representation of weights during computation, analogous to the multi-head attention mechanism. The rest of the workflow is straightforward: we query the expert label vectors within the semantic space. Experts with high semantic co-occurrence possess knowledge highly relevant to the query.</p>
<p>But here arises another problem: how can we ensure that these label vectors truly represent the capabilities of each expert? In other words, we need an inherent communication mechanism within the model to establish a correlation between the label vectors and the output capabilities of the experts. We can easily notice two simple and elegant methods that require no complex mapping and can associate the network's outputs with labels.</p>
<ol>
<li>Treat the label vector itself as the bias term of the expert network.</li>
<li>Map the label vector into part of the weights via low-rank factorization.</li>
</ol>
<p>In most implementations, the routing score is used as the fusion weight among multiple experts. We may regard the weighted summation of fused routing scores itself as another transformation operation of vectors in the embedding space. From this perspective, it is not difficult to realize that this is essentially a kind of implicit internal semantic transformation. This process is somewhat similar to the brain activity of &quot;association&quot; that humans often engage in. However, human association is attention-aware — in fact, human attention pervades the entire thinking process <a href="#ref-buschman2010goal">(Buschman &amp; Miller, 2010)</a>. This is a level that current AI can hardly reach, whether in terms of existing theories or computing hardware itself. We may require quantum-state computing to stack attention with different possibilities so as to achieve the goal of human-like association.</p>
<p>In essence, achieving interpretability for FFNs requires deriving their transformation dynamics from the representational manifold of the embedding space. We can roughly list two simple ways to use routing weight slicing as a bias term:</p>
<div class="math math--block" data-tex="X_t = S_e * ffn(X_{t-1})+B_e">X_t = S_e * ffn(X_{t-1})+B_e</div>
<div class="math math--block" data-tex="X_t = S_e * (ffn(X_{t-1}) + B_e)">X_t = S_e * (ffn(X_{t-1}) + B_e)</div>
<p>The difference between these two formulas lies in whether the gradient passes through the routing score. As can be clearly observed, their differential matrices differ only by an extra multiplication with the routing score. Comparative experiments demonstrate that this discrepancy is negligible. However, regardless of the type of bias adopted, its performance is consistently better than the bias-free counterpart.</p>
<p>In fact, we can also explore more complex mapping methods. This paper only takes the current MoE architecture as an example: using the sliced mapping weights of its routing scores as bias terms can better align the semantic transformation between inputs and outputs.</p>
<p><img src="assets/ffn_compare.webp" alt="FFN Compare" width="1761" height="893" srcset="assets/ffn_compare-820.webp 820w, assets/ffn_compare-1640.webp 1640w, assets/ffn_compare.webp 1761w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p><strong>Exact Linear Attention GPT</strong></p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center">Inner bias</th>
<th style="text-align:center">Outer bias</th>
<th style="text-align:center">Without bias</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/fgpt_exp+hadm_fib_loss.webp" alt="Inner bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/fgpt_exp+hadm_fob_loss.webp" alt="Outer bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/fgpt_exp+hadm_loss.webp" alt="Without bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<p><strong>Full Attention GPT</strong></p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center">Inner bias</th>
<th style="text-align:center">Outer bias</th>
<th style="text-align:center">Without bias</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/gpt_fib_loss.webp" alt="Inner bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/gpt_fob_loss.webp" alt="Outer bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/gpt_loss.webp" alt="Without bias" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<blockquote>
<p><strong>Figure:</strong> Comparison of Exact Linear Attention GPT (top row) and Full Attention GPT (bottom row).</p>
</blockquote>
<p>As for the issue of communication overhead, token dispatching based on routing scores is currently irreplaceable owing to the inherent nature of sparse activation. Nevertheless, cross-device transmission can be uniformly scheduled, analogous to the design of unified memory architecture <a href="#ref-jia2020megatron">(Jia et al., 2020)</a>. In fact, consecutive token blocks form semantic communities. <strong>Partitioning tokens into blocks in a proper manner can substantially reduce communication overhead, compared with fine-grained routing conducted at the individual token level.</strong></p>
<h3 id="5-2-hyperlink-based-residual-replacement-for-degradation-eradication">5.2 Hyperlink-Based Residual Replacement for Degradation Eradication</h3>
<p>Traditional residual connections across multiple Decoder layers suffer from gradient vanishing and difficult cross-layer information propagation. The current mainstream solutions to this issue are HC (Hyper-Connection) and mHC (Manifold-Constrained Hyper-Connection).</p>
<p>We propose to reconstruct the residual pathway itself: we establish residual connections between Decoder layers at different depths and remove the attention residual branch in the standard Pre-Norm architecture, treating the entire Transformer layer as an integrated whole. Furthermore, since modern FFNs are equipped with gated structures, the gated outputs can be naturally leveraged to adaptively modulate the signal of each layer.</p>
<p><img src="assets/hyperlink.webp" alt="Hyperlink" width="1774" height="887" srcset="assets/hyperlink-820.webp 820w, assets/hyperlink-1640.webp 1640w, assets/hyperlink.webp 1774w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p>Experimental results demonstrate that our method can effectively accelerate training speed and substantially mitigate gradient degradation. Under identical computational overhead, Hyper-Link achieves faster convergence and better fitting performance than conventional Residual-Link. This also explains why the convergence curves in all experimental plots show an extremely steep initial drop followed by steady decline in the later stage.</p>
<p><strong>In the experiment, we removed the final normalization layer of GPT to accelerate convergence speed</strong>. In practical cluster training, however, the final normalization is required to maintain model stability. For connections between hyperlinks under gated mechanisms, output normalization is unnecessary, as the gate itself modulates the output. In particular, for extremely large and diverse datasets containing tens of thousands of tokens with high semantic entropy in the corpus (spanning multiple domains), additional normalization (segmented normalization) is needed to sustain training stability.</p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center">Hyper-Link</th>
<th style="text-align:center">Normal</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/gpt_loss.webp" alt="Hyper-Link" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/std_gpt_loss.webp" alt="Normal" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<blockquote>
<p><strong>Figure:</strong> Training Comparison (GPT)</p>
</blockquote>
<h3 id="5-3-how-memory-works">5.3 How Memory Works</h3>
<p>In general, human memory exists in two forms. The first is what we term <strong>factual memory</strong>, which records that a certain event has occurred. The second is <strong>qualitative memory</strong>, which represents how a given event is perceived or evaluated. This fundamental dichotomy of memory divides all known information into two categories: behavioral judgment and objective existence.</p>
<p>For factual memory, we typically regard it as background knowledge. Qualitative memory, by contrast, functions more like inherent constraints and rules. A simple example illustrates this point: suppose you dine at a restaurant one day and have a poor experience. Would you choose to visit again? Evidently, it is the <strong>known judgment content</strong> embedded in qualitative memory that guides your subsequent decision-making.</p>
<p>In conventional model training, this cognitive capability is entirely encapsulated within the Feed-Forward Network (FFN), forming an inexplicable black box where multiple conditional constraints are tightly coupled together. If we aim to explicitly disentangle factual memory from qualitative memory, we need to redesign the entire computational process from the perspective of semantic space transformation.</p>
<p>According to recent research by the DeepSeek team, the Engram <a href="#ref-cheng2026conditional">(Cheng et al., 2026)</a> module performs remarkably well as an auxiliary component for knowledge storage, corresponding to factual memory. This raises a key question: how should we construct qualitative memory, which serves as the more critical behavioral guideline itself? Our answer to this is clearly: <strong>Attention is all you need.</strong></p>
<p>Do you still remember that we mentioned earlier in Hyper-Link that <strong>we removed the attention residual</strong>? This is not merely to make the layer output act as a whole; instead, we have another ingenious application for it here.</p>
<p>If we perform discrete differentiation on the process <span class="math math--inline" data-tex="X_{k} \to X_{k-1}">X_{k} \to X_{k-1}</span>, we can observe that:</p>
<div class="math math--block" data-tex="X_{k} = DecoderLayer(X_{k-1})">X_{k} = DecoderLayer(X_{k-1})</div>
<div class="math math--block" data-tex="\Delta X_{k|k-1} = X_{k} - X_{k-1} = ffn(attn(RMSnorm(X_{k-1})\,)\,)">\Delta X_{k|k-1} = X_{k} - X_{k-1} = ffn(attn(RMSnorm(X_{k-1})\,)\,)</div>
<p>We refer to the differential result <span class="math math--inline" data-tex="\Delta X_{k|k-1}">\Delta X_{k|k-1}</span> as the &quot;Flow&quot; of the transformation <span class="math math--inline" data-tex="X_{k} \to X_{k-1}">X_{k} \to X_{k-1}</span>. It serves as the &quot;trajectory&quot; of the semantic transformation process, i.e., an object that records how the semantics evolve after passing through the current layer. Then we design a bidirectional attention-based perception module for the &quot;Flow&quot; of this evolution process, which is formulated as follows:</p>
<ul>
<li><span class="math math--inline" data-tex="Q \in \mathbb{R}^{D \times D}">Q \in \mathbb{R}^{D \times D}</span> is flow's query representation, means &quot;What about this transformation.&quot;</li>
<li><span class="math math--inline" data-tex="K \in \mathbb{R}^{D \times D}">K \in \mathbb{R}^{D \times D}</span> is flow's key representation, means &quot;What I can provide for this transformation.&quot;</li>
<li><span class="math math--inline" data-tex="V \in \mathbb{R}^{D \times D}">V \in \mathbb{R}^{D \times D}</span> is flow's value representation, means &quot;What I can do for this transformation.&quot;</li>
</ul>
<p><strong>Pseudocode:</strong></p>
<figure class="codeblock" data-lang="python"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">python</span><span class="codeblock__lang">python</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-python hl"><span class="cl"><span class="k">def</span><span class="w"> </span><span class="nf">lob</span><span class="p">(</span><span class="n">dx</span><span class="p">):</span></span><span class="cl">    <span class="n">q</span> <span class="o">=</span> <span class="n">Q</span><span class="p">(</span><span class="n">dx</span><span class="p">)</span></span><span class="cl">    <span class="n">k</span> <span class="o">=</span> <span class="n">K</span><span class="p">(</span><span class="n">dx</span><span class="p">)</span></span><span class="cl">    <span class="n">v</span> <span class="o">=</span> <span class="n">V</span><span class="p">(</span><span class="n">dx</span><span class="p">)</span></span><span class="cl">    <span class="c1"># Bidirectional Linear Attention</span></span><span class="cl">    <span class="k">return</span> <span class="o">=</span> <span class="n">ELA</span><span class="p">(</span><span class="n">q</span><span class="p">,</span> <span class="n">k</span><span class="p">,</span> <span class="n">v</span><span class="p">)</span></span><span class="cl"><span class="o">...</span></span><span class="cl"><span class="k">def</span><span class="w"> </span><span class="nf">decoder</span><span class="p">(</span><span class="n">x</span><span class="p">):</span></span><span class="cl">    <span class="n">x_norm</span> <span class="o">=</span> <span class="n">norm</span><span class="p">(</span><span class="n">x</span><span class="p">)</span></span><span class="cl">    <span class="n">attn</span> <span class="o">=</span> <span class="n">ELA_causal</span><span class="p">(</span><span class="n">query</span><span class="o">=</span><span class="n">x_norm</span><span class="p">,</span></span><span class="cl">        <span class="n">key</span><span class="o">=</span><span class="n">x_norm</span><span class="p">,</span></span><span class="cl">        <span class="n">value</span><span class="o">=</span><span class="n">x_norm</span><span class="p">)</span></span><span class="cl">    <span class="n">ffn_out</span><span class="p">,</span> <span class="n">aux_loss</span> <span class="o">=</span> <span class="n">MoE</span><span class="p">(</span><span class="n">attn</span><span class="p">)</span></span><span class="cl">    <span class="c1"># get the flow query attention output</span></span><span class="cl">    <span class="n">lob_out</span> <span class="o">=</span> <span class="n">lob</span><span class="p">(</span><span class="n">ffn_out</span><span class="p">)</span></span><span class="cl">    <span class="c1"># hyper-link</span></span><span class="cl">    <span class="k">return</span> <span class="n">x</span> <span class="o">+</span> <span class="n">ffn_out</span> <span class="o">+</span> <span class="n">lob_out</span><span class="p">,</span> <span class="n">aux_loss</span></span><span class="cl"></span></code></pre></div></figure>
<blockquote>
<p><em>For detailed implementation, please refer to our GitHub repository.</em></p>
</blockquote>
<p>With the above construction, we observe that the DecoderLayer equipped with the Transformation Flow comprehensively outperforms the vanilla version in training. Datasets that originally required 30 epochs for convergence only need around 10 epochs after integrating the memory module, and the training loss and validation loss become much more consistent.</p>
<p>In fact, this is a mathematical formulation of <strong>qualitative memory</strong>. The QKV weight matrices of bidirectional attention can &quot;memorize&quot; which representations lead to lower loss during training. Our input is the <strong>layer-wise Transformation</strong> itself, enabling the model to implicitly record the layer's processing experience through learning to serve subsequent generation.</p>
<p>Since the output of the FFN is produced by causal attention, it inherently possesses forward causal properties. Meanwhile, we require memory to query the transformation history of all positions, making this a global bidirectional attention query process.</p>
<p>Although this training process appears to adopt explicit supervised learning, it essentially leverages supervised learning to implement reinforcement learning. This is because the entire pipeline relies on parameterized memory content queried from layer Transformations to support final output, forming an implicit reinforcement learning paradigm—the &quot;Action-Reward&quot; mechanism: the current memory query serves as the <strong>Action</strong>, and its direct contribution to the loss acts as the corresponding <strong>Reward</strong>.</p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center">Memory Lobe</th>
<th style="text-align:center">Normal</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/fgpt_mem_loss.webp" alt="Memory lobe" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/fgpt_exp+hadm_loss.webp" alt="Normal" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<blockquote>
<p><strong>Figure:</strong> Training Comparison (ELA GPT)</p>
</blockquote>
<p>Furthermore, additionally, the QKV weight matrices of this memory module are pluggable. Theoretically, this framework can be embedded into any semantic-transformation-based model that is capable of producing <span class="math math--inline" data-tex="\Delta X_{k|k-1}">\Delta X_{k|k-1}</span>, allowing it to learn internal experience and form qualitative memory. This provides a brand-new paradigm for LLM training beyond <strong>LoRA</strong> and <strong>Engram</strong> methods.</p>
<p>In particular, the design inspiration of this module is derived from the principle of biological neural memory <a href="#ref-polyn2008">(Polyn &amp; Kahana, 2008)</a> <a href="#ref-zhang2018">(Zhang et al., 2018)</a>, where the prefrontal cortex plays a crucial role in the contextual integration of memory <a href="#ref-desousa2026">(de Sousa et al., 2026)</a>.</p>
<hr />
<h2 id="6-experiments">6. Experiments</h2>
<p><strong>To unify the experimental variables, all attention kernels involved in the corresponding attention model adopt the same type.</strong></p>
<p>These two models are built for ablation validation. The training dataset contains 129×3500 samples, amounting to 451,500 tokens. We adopt the Minimind <a href="#ref-jingyao2026">(Gong, 2026)</a> tokenizer with a vocabulary size of V=6400. The model architecture adopts L=4 Transformer layers, with a model dimension dmodel=256 and nheads=4 attention heads. A Mixture-of-Experts (MoE) module is further introduced with nexperts=4. The total number of model parameters is 5,838,864. Training with 30 epochs.</p>
<p>We separately train FA-GPT with standard MoE, as well as ELA-GPT variants equipped with the Hadamard Exp Kernel and the Summation Squared Euclidean Distance Kernel.</p>
<h3 id="6-1-no-memory-training-comparison">6.1 No Memory Training Comparison</h3>
<p><img src="assets/Architecture.webp" alt="Architecture" width="1536" height="1024" srcset="assets/Architecture-820.webp 820w, assets/Architecture.webp 1536w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p><strong>Training Comparison (Hyper-Link)</strong></p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center"><span class="math math--inline" data-tex="|A_i+B_j|^2">|A_i+B_j|^2</span></th>
<th style="text-align:center"><span class="math math--inline" data-tex="\exp(A_i)\exp(B_j)">\exp(A_i)\exp(B_j)</span></th>
<th style="text-align:center">Full</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/fgpt_L2_loss.webp" alt="L2" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/fgpt_exp+hadm_loss.webp" alt="EH" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/gpt_loss.webp" alt="Full" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<p><strong>Training Comparison (Normal)</strong></p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center"><span class="math math--inline" data-tex="|A_i+B_j|^2">|A_i+B_j|^2</span></th>
<th style="text-align:center"><span class="math math--inline" data-tex="\exp(A_i)\exp(B_j)">\exp(A_i)\exp(B_j)</span></th>
<th style="text-align:center">Full</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/std_fgpt_L2_loss.webp" alt="L2" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/std_fgpt_eh_loss.webp" alt="EH" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/std_gpt_loss.webp" alt="Full" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<p>There it can be observed that the two models exhibit negligible differences in training performance. In particular, the ELA variant shows a slight advantage in anti-overfitting ability.</p>
<h3 id="6-2-memory-additional">6.2 Memory Additional</h3>
<p>In this comparative experiment, after integrating the Memory module, the model not only achieves faster loss convergence. On the vanilla GPT, we also observe an abrupt drop with an inflection point at around the 750th training step (counted as global steps, with 10 epochs totaling 1750 steps).</p>
<p>This is certainly not a sudden &quot;grokking&quot; of the model. Instead, the memory module comes into play and enables the model to capture underlying patterns. Given the relatively small scale of training tokens in our setup, the number of non-embedding parameters increases to 6,624,272 after introducing the memory module, allowing the model to directly reuse learned empirical regularities.</p>
<p>When a causal mask is applied to the Memory Query of the vanilla GPT, such abrupt performance drop vanishes completely.</p>
<p>These experimental results demonstrate that the proposed ELA exhibits excellent performance in anti-overfitting and generalization capability.</p>
<p><strong>Training Comparison (Hyper-Link &amp; Memory)</strong></p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:center"><span class="math math--inline" data-tex="\exp(A_i)\exp(B_j)">\exp(A_i)\exp(B_j)</span></th>
<th style="text-align:center">Full</th>
<th style="text-align:center">Full (mem-causal)</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:center"><img src="assets/fgpt_mem_loss.webp" alt="EH" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/gpt_mem_loss.webp" alt="Full" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
<td style="text-align:center"><img src="assets/gpt_mem_causal_loss.webp" alt="Causal" width="640" height="480" loading="lazy" decoding="async" class="post-image"></td>
</tr>
</tbody>
</table></div>
<h3 id="6-3-long-range-training">6.3 Long-range Training</h3>
<p>In long-range training, ELA maintains stable convergence.</p>
<p><img src="assets/fgpt_exp+hadmE50_loss.webp" alt="Long-range Training" width="640" height="480" loading="lazy" decoding="async" class="post-image"></p>
<hr />
<h2 id="7-extend-to-vision-models">7. Extend to Vision Models</h2>
<p>We reformulate deep convolutions in YOLO <a href="#ref-yolo26">(YOLO26)</a> <a href="#ref-yolov8">(YOLOv8)</a> with linear attention to build a model featuring fewer parameters and lower inference latency, which delivers outstanding performance on our benchmarks.</p>
<p><img src="assets/benchmark_plot.webp" alt="CUDA vs CPU" width="2234" height="654" srcset="assets/benchmark_plot-820.webp 820w, assets/benchmark_plot-1640.webp 1640w, assets/benchmark_plot.webp 2234w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p><img src="assets/benchmark_3way.webp" alt="YOLO-LAT vs YOLOv26 Speed" width="2224" height="1684" srcset="assets/benchmark_3way-820.webp 820w, assets/benchmark_3way-1640.webp 1640w, assets/benchmark_3way.webp 2224w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p><img src="assets/benchmark_accuracy.webp" alt="YOLO-LAT vs YOLOv26 Accuracy" width="2235" height="657" srcset="assets/benchmark_accuracy-820.webp 820w, assets/benchmark_accuracy-1640.webp 1640w, assets/benchmark_accuracy.webp 2235w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p>In comparative experiments, YOLO-LAT achieves <strong>2.2× faster inference on CPU</strong> and <strong>4.3× faster inference on GPU</strong> compared with vanilla YOLO <a href="#ref-yolo26">(YOLO26)</a>. In terms of model performance, our method obtains competitive accuracy with <strong>7.9× fewer parameters</strong>: YOLO-LAT reaches an mAP@0.5 of 0.962, close to YOLO's 0.998. Nevertheless, there exists a noticeable gap in mAP@0.5:0.95 (0.515 versus 0.951), indicating inferior bounding box localization precision.</p>
<p>We verify that this limitation stems from the lack of depth information. Traditional YOLO adopts CASC dynamic channel pruning <a href="#ref-yolov5">(YOLOv5)</a> to simulate hierarchical visual perception. In contrast, YOLO-LAT leverages inherent attention mechanisms to focus on foreground objects without dedicated modules tailored for object detection. Such results sufficiently demonstrate the effectiveness of generalizing linear attention <a href="#ref-linearvit">(LinearViT)</a> to visual models.</p>
<p>To further elaborate on the effects of Exact Linear Attention on Vision Transformer (ViT) models, we constructed a simplified object detection model based on FCOS <a href="#ref-tian2019fcos">(Tian et al., 2019)</a>. Given its excellent training performance, we adopted the results from 50 basic training epochs to prevent overfitting. In fact, the linear-complexity attention computation of Exact Linear Attention is particularly well-suited for pure vision models. Such models can natively derive the corresponding Q, K and V matrices via convolution operations and subsequently perform attention calculations on the entire image, which aligns with the human visual mechanism of focusing on specific targets.</p>
<p><img src="assets/detection_viz.webp" alt="Detection Attention Vision" width="2967" height="1478" srcset="assets/detection_viz-820.webp 820w, assets/detection_viz-1640.webp 1640w, assets/detection_viz.webp 2967w" sizes="(max-width: 900px) 100vw, 1280px" loading="lazy" decoding="async" class="post-image"></p>
<p>In future work, we will introduce vision-lidar fused images with native depth cues to train models equipped with inherent depth estimation capabilities. This paper preliminarily explores future development trends of visual models, and proposes a technical paradigm that compensates for the world understanding defects of world models via depth information estimation.</p>
<hr />
<h2 id="8-exploratory-works">8. Exploratory Works</h2>
<p>Subsequent work will further explore extended designs based on this paper.</p>
<ul>
<li><strong>Generalization to Diffusion Models</strong>: leveraging infinitely long precise attention to globally perceive all fine-grained details and thereby boost generation quality.</li>
<li>Explore decomposable kernel functions with properties closer to <span class="math math--inline" data-tex="e^{xy}">e^{xy}</span>. (The current form of the kernel function is <span class="math math--inline" data-tex="e^{x+y}">e^{x+y}</span>.)</li>
<li>Explore whether the memory module can break the limitations of scaling laws, enabling the model to achieve stronger performance with fewer parameters (e.g., by low-rank factorization of QK).</li>
<li>Extend to full-modality world models.</li>
</ul>
<p>All these endeavors are inseparable from our core insight: Exact Linear Attention.</p>
<hr />
<h2 id="acknowledgments">Acknowledgments</h2>
<blockquote>
<p>I sincerely appreciate the anonymous reviewers and the associate editor for their valuable time, rigorous reviews, and insightful constructive comments. Their professional feedback and thoughtful suggestions have greatly helped refine the technical presentation, consolidate the logical framework, and substantially improve the overall quality of this manuscript.</p>
</blockquote>
<p>作者感谢所有给予过帮助的朋友（排名不分先后）</p>
<ul>
<li><a href="https://github.com/mzwing" target="_blank" rel="noopener noreferrer">mzwing (Lockinwize Lolite)</a></li>
<li><a href="https://github.com/hibays" target="_blank" rel="noopener noreferrer">hibays (hibays)</a></li>
<li>.....</li>
</ul>
<hr />
<h2 id="references">References</h2>
<ol>
<li>
<p><a id="ref-katharopoulos2020"></a>A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, &quot;Transformers are RNNs: Fast autoregressive transformers with linear attention,&quot; in <em>Proc. 37th Int. Conf. Mach. Learn. (ICML)</em>, 2020, pp. 5156–5165.</p>
</li>
<li>
<p><a id="ref-schlag2021"></a>I. Schlag, K. Irie, and J. Schmidhuber, &quot;Linear transformers are secretly fast weight programmers,&quot; in <em>Proc. 38th Int. Conf. Mach. Learn. (ICML)</em>, 2021, pp. 9355–9366.</p>
</li>
<li>
<p><a id="ref-qin2022"></a>Z. Qin, X. Han, W. Sun, D. Li, L. Kong, N. Barnes, and Y. Zhong, &quot;The devil in linear transformer,&quot; in <em>Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP)</em>, 2022, pp. 7025–7041.</p>
</li>
<li>
<p><a id="ref-vaswani2017"></a>A. Vaswani et al., &quot;Attention is all you need,&quot; in <em>Proc. 31st Conf. Neural Inf. Process. Syst. (NeurIPS)</em>, 2017, pp. 5998–6008.</p>
</li>
<li>
<p><a id="ref-jingyao2026"></a>Jingyao Gong, &quot;MiniMind: Train a Tiny LLM from Scratch,&quot; GitHub: <a href="https://github.com/jingyaogong/minimind" target="_blank" rel="noopener noreferrer">https://github.com/jingyaogong/minimind</a></p>
</li>
<li>
<p><a id="ref-cheng2026conditional"></a>Xin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, et al., &quot;Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models,&quot; <em>arXiv preprint arXiv:2601.07372</em>, 2026.</p>
</li>
<li>
<p><a id="ref-zhu2024hyper"></a>Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, and Xun Zhou, &quot;Hyper-Connections,&quot; <em>arXiv preprint arXiv:2409.19606</em>, 2024.</p>
</li>
<li>
<p><a id="ref-xie2025mhc"></a>Zhenda Xie, Wentao Zhang, Xinyu Zhao, Yukai Li, Peng Wang, Weiran You, and others, &quot;mHC: Manifold-Constrained Hyper-Connections,&quot; <em>arXiv preprint arXiv:2512.24880</em>, 2025.</p>
</li>
<li>
<p><a id="ref-mercer1909"></a>J. Mercer, &quot;Functions of positive and negative type, and their connection with the theory of integral equations,&quot; <em>Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character</em>, vol. 209, pp. 415–446, 1909.</p>
</li>
<li>
<p><a id="ref-minimax2025lightning"></a>MiniMax Team, &quot;MiniMax-01: Scaling Foundation Models with Lightning Attention,&quot; <em>arXiv preprint arXiv:2501.08313</em>, 2025.</p>
</li>
<li>
<p><a id="ref-kimilinear2025"></a>Kimi Team, &quot;Kimi Linear: A Novel Hybrid Linear Attention Architecture,&quot; <em>arXiv preprint arXiv:2510.xxxxx</em>, 2025.</p>
</li>
<li>
<p><a id="ref-jain2019attention"></a>S. Jain and B. C. Wallace, &quot;Attention is not Explanation,&quot; in <em>Proc. Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT)</em>, 2019.</p>
</li>
<li>
<p><a id="ref-jia2020megatron"></a>Z. Jia, W. Kwon, and O. Ruwase, &quot;Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM,&quot; in <em>Proc. Int. Conf. High Performance Computing, Networking, Storage and Analysis (SC'20)</em>, 2020.</p>
</li>
<li>
<p><a id="ref-buschman2010goal"></a>T. J. Buschman and E. K. Miller, &quot;Goal-direction and top-down control,&quot; <em>Philosophical Transactions of the Royal Society B: Biological Sciences</em>, vol. 365, no. 1544, pp. 1271–1278, 2010.</p>
</li>
<li>
<p><a id="ref-polyn2008"></a>Polyn, S. M., &amp; Kahana, M. J., &quot;Memory search and the neural representation of context,&quot; <em>Trends in Cognitive Sciences</em>, 12(1):24–30, 2008.</p>
</li>
<li>
<p><a id="ref-zhang2018"></a>Zhang, W., van Ast, V. A., Klumpers, F., Roelofs, K., &amp; Hermans, E. J., &quot;Memory contextualization: The role of the left inferior frontal gyrus in binding event and contextual information,&quot; <em>Journal of Cognitive Neuroscience</em>, 30(5):698–713, 2018.</p>
</li>
<li>
<p><a id="ref-desousa2026"></a>de Sousa, A. F., Zeidler, Z. E., Almeida-Filho, D. G., Shen, Y., Luchetti, A., Simanian, S., Mardini, M., DeNardo, L. A., &amp; Silva, A. J., &quot;The prefrontal cortex controls memory organization in the hippocampus,&quot; <em>Nature Neuroscience</em>, 29:1191–1202, 2026. doi: 10.1038/s41593-026-02231-1.</p>
</li>
<li>
<p><a id="ref-yolov1"></a>J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, &quot;You only look once: Unified, real-time object detection,&quot; in <em>Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</em>, 2016, pp. 779–788.</p>
</li>
<li>
<p><a id="ref-yolov5"></a>Ultralytics, &quot;YOLOv5: A state-of-the-art real-time object detection system,&quot; 2020. [Online]. Available: <a href="https://github.com/ultralytics/yolov5" target="_blank" rel="noopener noreferrer">https://github.com/ultralytics/yolov5</a></p>
</li>
<li>
<p><a id="ref-yolov7"></a>C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, &quot;YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,&quot; in <em>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</em>, 2023, pp. 7464–7475.</p>
</li>
<li>
<p><a id="ref-yolov8"></a>Ultralytics, &quot;YOLOv8: A new state-of-the-art computer vision model,&quot; 2023. [Online]. Available: <a href="https://github.com/ultralytics/ultralytics" target="_blank" rel="noopener noreferrer">https://github.com/ultralytics/ultralytics</a></p>
</li>
<li>
<p><a id="ref-yolo12"></a>Ultralytics, &quot;YOLO12: Attention-centric object detection,&quot; 2025. [Online]. Available: <a href="https://docs.ultralytics.com/models/yolo12/" target="_blank" rel="noopener noreferrer">https://docs.ultralytics.com/models/yolo12/</a></p>
</li>
<li>
<p><a id="ref-yolo-dma"></a>Y. Li, Z. Wang, and H. Liu, &quot;YOLO-DMA: A small-object detector based on multi-scale deformable convolution and linear attention,&quot; <em>Electronics</em>, vol. 15, no. 4, p. 812, 2026.</p>
</li>
<li>
<p><a id="ref-yolo26"></a>G. Jocher and J. Qiu, &quot;Ultralytics YOLO26,&quot; version 26.0.0, 2026. [Online]. Available: <a href="https://github.com/ultralytics/ultralytics" target="_blank" rel="noopener noreferrer">https://github.com/ultralytics/ultralytics</a></p>
</li>
<li>
<p><a id="ref-linearvit"></a>C. Zheng, &quot;The linear attention resurrection in vision transformer,&quot; <em>arXiv preprint arXiv:2501.16182</em>, 2025.</p>
</li>
<li>
<p><a id="ref-tian2019fcos"></a>Z. Tian, C. Shen, H. Chen, and T. He, &quot;FCOS: Fully convolutional one-stage object detection,&quot; in <em>Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV)</em>, 2019, pp. 9627–9636, arXiv:1904.01355.</p>
</li>
</ol>
]]></content></entry>
<entry><title>你好，墨痕 —— 一个用 Python 渲染的水墨博客引擎</title><link href="https://yaunt.github.io/blog/blogs/hello-inkstone/app.html"/><id>https://yaunt.github.io/blog/blogs/hello-inkstone/app.html</id><updated>2026-10-10T00:00:00Z</updated><published>2026-07-05T00:00:00Z</published><author><name>曦儿</name></author><category term="静态站点"/><category term="Python"/><category term="构建系统"/><category term="前端"/><category term="设计"/><category term="Markdown"/><category term="渲染管线"/><category term="KaTeX"/><category term="Mermaid"/><category term="CoT"/><category term="思维链"/><category term="交互设计"/><summary type="text">这个站点由 Inkstone 渲染 —— 一个用 Python 写的静态博客引擎。它只做三件事：读 Markdown、读配置、写出一坨可以直接丢到任何静态托管上的文件。 设计上我给自己定了三条约束： 零运行时依赖：产物是纯 HTML / CSS…</summary><content type="html"><![CDATA[<section class="cot" data-cot><div class="cot__head"><button class="cot__toggle" type="button" aria-expanded="true"><span class="cot__icon" aria-hidden="true">◈</span><span class="cot__title">为什么要把推理过程单独拎出来</span><span class="cot__badge">2 步</span><span class="cot__chevron" aria-hidden="true"></span></button><div class="cot__actions"><button type="button" class="cot__action" data-cot-copy>复制</button><button type="button" class="cot__action" data-cot-export>导出</button><button type="button" class="cot__action" data-cot-mindmap>导图</button><button type="button" class="cot__action" data-cot-hide>隐藏</button></div></div><div class="cot__panel"><ol class="cot__steps"><li class="cot__step" data-step="1" data-title=""><span class="cot__index" aria-hidden="true">01</span><div class="cot__step-body"><div class="cot__step-content"><p>正文负责结论，推理过程负责可信度。</p>
</div></div></li><li class="cot__step" data-step="2" data-title=""><span class="cot__index" aria-hidden="true">02</span><div class="cot__step-body"><div class="cot__step-content"><p>把两者混在一起，读者要么被细节淹没，要么无法验证结论。分开之后，
想深究的人可以展开，只想看结论的人可以跳过。</p>
</div></div></li></ol></div></section><p>这个站点由 <strong>Inkstone</strong> 渲染 —— 一个用 Python 写的静态博客引擎。它只做三件事：读 Markdown、读配置、写出一坨可以直接丢到任何静态托管上的文件。</p>
<p>设计上我给自己定了三条约束：</p>
<ol>
<li><strong>零运行时依赖</strong>：产物是纯 HTML / CSS / JS，不依赖任何 CDN 才能跑起来；</li>
<li><strong>可移植</strong>：所有站内链接都是相对路径，放在 <code>/</code> 或 <code>/blog/</code> 或任何子目录都能直接用；</li>
<li><strong>可增量</strong>：重复构建时，未改动的文章直接复用缓存，不再走一遍渲染管线。</li>
</ol>
<p>这篇文章分两部分：先讲引擎本身（构建管线、视觉语言、实现细节），再把它能渲染的东西 —— Markdown 的全部语法与「思维链」区块 —— 逐项跑一遍。既是介绍，也是一份可以对着抄的能力清单。</p>
<h2 id="构建管线">构建管线</h2>
<p>从按下回车到产物落盘，一共十步。每一步都是独立的模块，可以单独替换：</p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:left">步骤</th>
<th style="text-align:left">模块</th>
<th style="text-align:left">职责</th>
<th style="text-align:left">增量策略</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left">1</td>
<td style="text-align:left"><code>engine/config.py</code></td>
<td style="text-align:left">配置合并、环境变量覆盖、校验</td>
<td style="text-align:left">—</td>
</tr>
<tr>
<td style="text-align:left">2</td>
<td style="text-align:left"><code>engine/theme.py</code></td>
<td style="text-align:left">主题加载与元数据</td>
<td style="text-align:left">—</td>
</tr>
<tr>
<td style="text-align:left">3</td>
<td style="text-align:left"><code>engine/content.py</code></td>
<td style="text-align:left">Frontmatter 解析、文章图谱</td>
<td style="text-align:left">文件 mtime + 内容哈希</td>
</tr>
<tr>
<td style="text-align:left">4</td>
<td style="text-align:left"><code>engine/assets.py</code></td>
<td style="text-align:left">WebP 转换、响应式变体</td>
<td style="text-align:left">源文件哈希</td>
</tr>
<tr>
<td style="text-align:left">5</td>
<td style="text-align:left"><code>engine/markdown.py</code></td>
<td style="text-align:left">Markdown → HTML 渲染管线</td>
<td style="text-align:left">文章签名 + 配置签名</td>
</tr>
<tr>
<td style="text-align:left">6</td>
<td style="text-align:left"><code>engine/search.py</code></td>
<td style="text-align:left">搜索索引与分片</td>
<td style="text-align:left">每次重建（成本极低）</td>
</tr>
<tr>
<td style="text-align:left">7</td>
<td style="text-align:left"><code>engine/builder.py</code></td>
<td style="text-align:left">静态资源压缩 + 指纹</td>
<td style="text-align:left">内容哈希</td>
</tr>
<tr>
<td style="text-align:left">8</td>
<td style="text-align:left"><code>engine/builder.py</code></td>
<td style="text-align:left">页面渲染</td>
<td style="text-align:left">复用第 5 步产物</td>
</tr>
<tr>
<td style="text-align:left">9</td>
<td style="text-align:left"><code>engine/feed.py</code></td>
<td style="text-align:left">sitemap / RSS / Atom</td>
<td style="text-align:left">每次重建</td>
</tr>
<tr>
<td style="text-align:left">10</td>
<td style="text-align:left"><code>engine/plugins.py</code></td>
<td style="text-align:left">插件钩子与收尾</td>
<td style="text-align:left">—</td>
</tr>
</tbody>
</table></div>
<blockquote>
<p>缓存的键是「文章内容哈希 + 资源映射哈希 + 相关配置哈希」。任何一项变了，这篇文章就会重新渲染；否则直接读上一次的 HTML 产物。</p>
</blockquote>
<aside class="admonition admonition--note"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">❋</span><span class="admonition__title">关于增量的边界</span></div><div class="admonition__body">
<p>增量构建只跳过<strong>逐篇文章</strong>的渲染。索引、sitemap、RSS 这些全局产物每次都重新生成 —— 它们的成本是毫秒级，做成增量反而增加出错概率。</p>
</div></aside>
<h2 id="视觉语言">视觉语言</h2>
<p>主题叫 <code>ink</code>，主色只有黑白两色，点缀交给「鎏银」：</p>
<ul>
<li><strong>纸</strong>：中性底色上叠一层 SVG 湍流噪声（<code>feTurbulence</code>），暗色模式下换成 <code>screen</code> 混合，让页面带上一点纸面的颗粒感。</li>
<li><strong>墨</strong>：Canvas 2D 画若干团缓慢漂移的径向渐变，作为背景的「墨云」；鼠标点击时墨点扩散成涟漪。</li>
<li><strong>银</strong>：卡片、面板与正文图片用 <code>mask-composite: exclude</code> 做出 1px 鎏银渐变描边，头像环与图片描边共用同一套金属渐变，悬浮时有一条高光扫过。</li>
<li><strong>留白</strong>：正文默认全宽（阅读设置里可随时收窄），文章页把目录放在左侧悬浮栏、随滚动高亮；行距 1.7，中文排版不缩进，靠段间距呼吸。</li>
</ul>
<p>所有视觉参数都在 <code>config.yaml</code> 与主题的 <code>theme.yaml</code> 里，改完重新构建即可；粒子密度、墨团数量、动效速度都能调。</p>
<h2 id="markdown-能力全景">Markdown 能力全景</h2>
<p>下面把引擎支持的 Markdown 语法逐项跑一遍，方便写新文章时对着抄。</p>
<h3 id="标题与目录">标题与目录</h3>
<p>从二级标题开始会被收进左侧目录，最多到四级。标题的锚点会保留中文，<code>### 标题与目录</code> 的锚点就是 <code>#标题与目录</code>。</p>
<h4 id="三级标题会缩进显示">三级标题会缩进显示</h4>
<h5 id="四级标题再缩进一层">四级标题再缩进一层</h5>
<h6 id="五级标题不进目录">五级标题不进目录</h6>
<h3 id="文本样式">文本样式</h3>
<p><strong>加粗</strong>、<em>斜体</em>、<s>删除线</s>、<code>行内代码</code>、<a href="https://example.com" target="_blank" rel="noopener noreferrer">外链</a>、<a href="../../blogs/exact-linear-attention/app.html" title="wikilink">站内链接</a>。</p>
<p>行内公式：质能方程 <span class="math math--inline" data-tex="E = mc^2">E = mc^2</span>，以及范数 <span class="math math--inline" data-tex="\|A_i + B_j\|^2">\|A_i + B_j\|^2</span>。</p>
<h3 id="块级公式">块级公式</h3>
<div class="math math--block" data-tex="\frac{\sum_{j=1}^{L} k(A_i, B_j)V_j}{\sum_{j=1}^{L} k(A_i, B_j)}
= \frac{\phi(A_i)\left[\sum_{j=1}^{L}\psi(B_j)^\top V_j\right]}{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top}">\frac{\sum_{j=1}^{L} k(A_i, B_j)V_j}{\sum_{j=1}^{L} k(A_i, B_j)}
= \frac{\phi(A_i)\left[\sum_{j=1}^{L}\psi(B_j)^\top V_j\right]}{\phi(A_i)\sum_{j=1}^{L}\psi(B_j)^\top}</div>
<p>也可以写成单行形式：$<span class="math math--inline" data-tex="k(A_i,B_j) = \exp(A_i) \ast \exp(B_j)">k(A_i,B_j) = \exp(A_i) \ast \exp(B_j)</span>$</p>
<h3 id="列表">列表</h3>
<p>无序列表：</p>
<ul>
<li>第一项</li>
<li>第二项
<ul>
<li>嵌套项</li>
<li>另一个嵌套项</li>
</ul>
</li>
<li>第三项</li>
</ul>
<p>有序列表：</p>
<ol>
<li>定位问题</li>
<li>提出假设</li>
<li>设计对照实验</li>
<li>收敛结论</li>
</ol>
<p>任务列表：</p>
<ul class="contains-task-list">
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> 渲染管线</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> 静态索引</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox"  type="checkbox"> AVIF 输出</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox"  type="checkbox"> 多语言互链</li>
</ul>
<h3 id="表格">表格</h3>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:left">内核函数</th>
<th style="text-align:left">分解形式</th>
<th style="text-align:left">语义</th>
<th style="text-align:center">计算复杂度</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left">加法平方欧氏距离</td>
<td style="text-align:left"><span class="math math--inline" data-tex="|A|^2 + |B|^2 + 2A\!\cdot\!B">|A|^2 + |B|^2 + 2A\!\cdot\!B</span></td>
<td style="text-align:left">同向证据检索</td>
<td style="text-align:center"><span class="math math--inline" data-tex="O(L)">O(L)</span></td>
</tr>
<tr>
<td style="text-align:left">减法平方欧氏距离</td>
<td style="text-align:left"><span class="math math--inline" data-tex="|A|^2 + |B|^2 - 2A\!\cdot\!B">|A|^2 + |B|^2 - 2A\!\cdot\!B</span></td>
<td style="text-align:left">对抗/矛盾检测</td>
<td style="text-align:center"><span class="math math--inline" data-tex="O(L)">O(L)</span></td>
</tr>
<tr>
<td style="text-align:left">Hadamard Exp</td>
<td style="text-align:left"><span class="math math--inline" data-tex="\exp(A) \ast \exp(B)">\exp(A) \ast \exp(B)</span></td>
<td style="text-align:left">特征共激活</td>
<td style="text-align:center"><span class="math math--inline" data-tex="O(L)">O(L)</span></td>
</tr>
<tr>
<td style="text-align:left">Softmax（标准注意力）</td>
<td style="text-align:left">不可精确分解</td>
<td style="text-align:left">方向相似度</td>
<td style="text-align:center"><span class="math math--inline" data-tex="O(L^2)">O(L^2)</span></td>
</tr>
</tbody>
</table></div>
<h3 id="代码块">代码块</h3>
<p>带文件名与行号：</p>
<figure class="codeblock codeblock--linenos" data-lang="python"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">engine/search.py</span><span class="codeblock__lang">python</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-python hl"><span class="cl"><span class="k">def</span><span class="w"> </span><span class="nf">build_search_payloads</span><span class="p">(</span><span class="n">context</span><span class="p">:</span> <span class="n">SiteContext</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="nb">dict</span><span class="p">[</span><span class="nb">str</span><span class="p">,</span> <span class="n">Any</span><span class="p">]:</span></span><span class="cl"><span class="w">    </span><span class="sd">"""生成索引数据（不落盘），便于测试与复用。"""</span></span><span class="cl">    <span class="n">shard_size</span> <span class="o">=</span> <span class="nb">max</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="nb">int</span><span class="p">(</span><span class="n">config</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s2">"shard_size"</span><span class="p">,</span> <span class="mi">100</span><span class="p">)))</span></span><span class="cl">    <span class="n">entries</span><span class="p">,</span> <span class="n">shards</span><span class="p">,</span> <span class="n">current</span> <span class="o">=</span> <span class="p">[],</span> <span class="p">[],</span> <span class="p">[]</span></span><span class="cl">    <span class="k">for</span> <span class="n">post</span> <span class="ow">in</span> <span class="n">context</span><span class="o">.</span><span class="n">posts</span><span class="p">:</span></span><span class="cl">        <span class="n">current</span><span class="o">.</span><span class="n">append</span><span class="p">({</span><span class="s2">"id"</span><span class="p">:</span> <span class="n">post</span><span class="o">.</span><span class="n">id</span><span class="p">,</span> <span class="s2">"content"</span><span class="p">:</span> <span class="n">strip_markdown</span><span class="p">(</span><span class="n">post</span><span class="o">.</span><span class="n">raw_md</span><span class="p">)})</span></span><span class="cl">        <span class="k">if</span> <span class="nb">len</span><span class="p">(</span><span class="n">current</span><span class="p">)</span> <span class="o">&gt;=</span> <span class="n">shard_size</span><span class="p">:</span></span><span class="cl">            <span class="n">shards</span><span class="o">.</span><span class="n">append</span><span class="p">({</span><span class="s2">"entries"</span><span class="p">:</span> <span class="n">current</span><span class="p">})</span></span><span class="cl">            <span class="n">current</span> <span class="o">=</span> <span class="p">[]</span></span><span class="cl">    <span class="k">return</span> <span class="p">{</span><span class="s2">"index"</span><span class="p">:</span> <span class="n">entries</span><span class="p">,</span> <span class="s2">"shards"</span><span class="p">:</span> <span class="n">shards</span><span class="p">}</span></span><span class="cl"></span></code></pre></div></figure>
<p>可折叠的长代码块（点击标题栏展开/收起）：</p>
<figure class="codeblock codeblock--collapsible" data-lang="bash" data-collapsed="true"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">deploy-blog.yml</span><span class="codeblock__lang">bash</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-bash hl"><span class="cl">name:<span class="w"> </span>Build<span class="w"> </span>and<span class="w"> </span>Deploy<span class="w"> </span>Blog</span><span class="cl">on:</span><span class="cl"><span class="w">  </span>push:</span><span class="cl"><span class="w">    </span>branches:<span class="w"> </span><span class="o">[</span>main<span class="o">]</span></span><span class="cl">jobs:</span><span class="cl"><span class="w">  </span>build:</span><span class="cl"><span class="w">    </span>runs-on:<span class="w"> </span>ubuntu-latest</span><span class="cl"><span class="w">    </span>steps:</span><span class="cl"><span class="w">      </span>-<span class="w"> </span>uses:<span class="w"> </span>actions/checkout@v4</span><span class="cl"><span class="w">      </span>-<span class="w"> </span>uses:<span class="w"> </span>actions/setup-python@v5</span><span class="cl"><span class="w">        </span>with:</span><span class="cl"><span class="w">          </span>python-version:<span class="w"> </span><span class="s1">'3.12'</span></span><span class="cl"><span class="w">      </span>-<span class="w"> </span>run:<span class="w"> </span>pip<span class="w"> </span>install<span class="w"> </span>-r<span class="w"> </span>requirements.txt</span><span class="cl"><span class="w">      </span>-<span class="w"> </span>run:<span class="w"> </span>python<span class="w"> </span>build.py</span><span class="cl"></span></code></pre></div></figure>
<p>其他语言：</p>
<figure class="codeblock" data-lang="javascript"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">app.js</span><span class="codeblock__lang">javascript</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-javascript hl"><span class="cl"><span class="kd">const</span><span class="w"> </span><span class="nx">state</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nx">running</span><span class="o">:</span><span class="w"> </span><span class="kc">false</span><span class="p">,</span><span class="w"> </span><span class="nx">raf</span><span class="o">:</span><span class="w"> </span><span class="mf">0</span><span class="w"> </span><span class="p">};</span></span><span class="cl"><span class="kd">function</span><span class="w"> </span><span class="nx">loop</span><span class="p">(</span><span class="nx">time</span><span class="p">)</span><span class="w"> </span><span class="p">{</span></span><span class="cl"><span class="w">  </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="o">!</span><span class="nx">state</span><span class="p">.</span><span class="nx">running</span><span class="p">)</span><span class="w"> </span><span class="k">return</span><span class="p">;</span></span><span class="cl"><span class="w">  </span><span class="nx">draw</span><span class="p">(</span><span class="nx">time</span><span class="p">);</span></span><span class="cl"><span class="w">  </span><span class="nx">state</span><span class="p">.</span><span class="nx">raf</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nx">requestAnimationFrame</span><span class="p">(</span><span class="nx">loop</span><span class="p">);</span></span><span class="cl"><span class="p">}</span></span><span class="cl"></span></code></pre></div></figure>
<figure class="codeblock" data-lang="yaml"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">config.yaml</span><span class="codeblock__lang">yaml</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-yaml hl"><span class="cl"><span class="nt">site</span><span class="p">:</span></span><span class="cl"><span class="w">  </span><span class="nt">title</span><span class="p">:</span><span class="w"> </span><span class="l l-Scalar l-Scalar-Plain">墨痕</span></span><span class="cl"><span class="w">  </span><span class="nt">url</span><span class="p">:</span><span class="w"> </span><span class="l l-Scalar l-Scalar-Plain">https://example.com</span></span><span class="cl"></span></code></pre></div></figure>
<p>无语言标注的代码块会按纯文本处理。</p>
<h3 id="引用与提示块">引用与提示块</h3>
<blockquote>
<p>真正困难的部分不是把 <span class="math math--inline" data-tex="O(L^2)">O(L^2)</span> 写成 <span class="math math--inline" data-tex="O(L)">O(L)</span>，而是证明这样写出来的东西<strong>仍然是原来那个东西</strong>。</p>
</blockquote>
<aside class="admonition admonition--note"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">❋</span><span class="admonition__title">说明</span></div><div class="admonition__body">
<p>普通说明信息，用于补充上下文。</p>
</div></aside>
<aside class="admonition admonition--tip"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">✦</span><span class="admonition__title">技巧</span></div><div class="admonition__body">
<p>按 <kbd>Ctrl</kbd> + <kbd>K</kbd> 可以随时唤起搜索面板；在文章页按 <kbd>J</kbd> / <kbd>K</kbd> 切换上一篇 / 下一篇。</p>
</div></aside>
<aside class="admonition admonition--important"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">❖</span><span class="admonition__title">重要</span></div><div class="admonition__body">
<p>重要提醒用鎏银描边，视觉上更醒目。</p>
</div></aside>
<aside class="admonition admonition--warning"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">⚠</span><span class="admonition__title">注意</span></div><div class="admonition__body">
<p>这条路径在 Safari 上需要 <code>-webkit-mask-composite</code>，标准写法是 <code>mask-composite: exclude</code>。</p>
</div></aside>
<aside class="admonition admonition--danger"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">⛔</span><span class="admonition__title">危险</span></div><div class="admonition__body">
<p>不要在生产环境直接 <code>rm -rf output/</code>，构建脚本已经会处理清理逻辑。</p>
</div></aside>
<aside class="admonition admonition--quote"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">❝</span><span class="admonition__title">引用</span></div><div class="admonition__body">
<p>好的工程不是把复杂度消灭，而是把它搬到不会伤人的地方。</p>
</div></aside>
<h3 id="脚注">脚注</h3>
<p>线性注意力最早的系统性讨论可以追溯到 Katharopoulos 等人的工作<sup class="footnote-ref"><a href="#fn1" id="fnref1">[1]</a></sup>，以及 Schlag 等人的「快速权重编程器」视角<sup class="footnote-ref"><a href="#fn2" id="fnref2">[2]</a></sup>。</p>
<h3 id="定义列表">定义列表</h3>
<dl>
<dt>增量构建</dt>
<dd>只重新渲染输入发生变化的文章，其余复用上一次的产物。</dd>
<dt>资源指纹</dt>
<dd>按内容哈希重命名 CSS / JS 文件，使浏览器可以长期缓存它们。</dd>
</dl>
<h3 id="图表">图表</h3>
<div class="mermaid" data-src-hash="24523932">
flowchart LR
    A[Markdown] --&gt; B{Frontmatter 解析}
    B --&gt; C[建立文章图谱]
    C --&gt; D[处理文章资源]
    D --&gt; E[渲染正文]
    E --&gt; F[生成搜索索引]
    C --&gt; G[渲染页面]
    F --&gt; G
    G --&gt; H[output/]
    H --&gt; I[Service Worker 预缓存]

</div>
<div class="mermaid" data-src-hash="60624703">
sequenceDiagram
    participant U as 访客
    participant W as Web Worker
    participant C as Cache API
    U-&gt;&gt;W: 输入关键词
    W-&gt;&gt;C: 读取分片索引
    C--&gt;&gt;W: blog-list-1.json
    W--&gt;&gt;U: 排序后的结果 + 高亮片段

</div>
<h3 id="图片">图片</h3>
<p>覆盖文章的资源目录 <code>assets/</code>，正文里用相对路径引用即可，构建时自动转 WebP、生成响应式变体并懒加载：</p>
<figure class="codeblock" data-lang="markdown"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">markdown</span><span class="codeblock__lang">markdown</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-markdown hl"><span class="cl">![<span class="nt">架构示意</span>](<span class="na">assets/architecture.png</span>)</span><span class="cl"></span></code></pre></div></figure>
<h3 id="双向链接">双向链接</h3>
<p>用 <code>[[文章标题]]</code> 或 <code>[[文章标题|显示文字]]</code> 建立链接，被引用的文章页面底部会自动出现「被引用」区块。比如指向 <a href="../../blogs/exact-linear-attention/app.html" title="wikilink">Exact Linear Attention</a>。</p>
<p>如果目标不存在，会渲染成一个带删除线的失效链接，方便在写作阶段发现断链。</p>
<h3 id="分隔线">分隔线</h3>
<hr />
<h2 id="思维链-让推理过程可以被查看-也可以被折叠">思维链：让推理过程可以被查看，也可以被折叠</h2>
<p>写技术文章时，最难的往往不是「说清楚结论」，而是「让结论可被检验」。结论写在正文里，推理过程折叠进 CoT 区块，是我目前找到的比较好的平衡。</p>
<p>引擎提供两种 CoT 写法。</p>
<h3 id="写法一-正文内联容器">写法一：正文内联容器</h3>
<p>在需要的位置插入 <code>:::cot</code> 容器，容器内用三级标题切分步骤：</p>
<figure class="codeblock" data-lang="markdown"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">markdown</span><span class="codeblock__lang">markdown</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-markdown hl"><span class="cl">:::cot 为什么选择 Hadamard Exp 核</span><span class="cl"><span class="gu">### 观察</span></span><span class="cl">候选核函数里，只有指数型同时满足非负性与精确可分解。</span><span class="cl"></span><span class="cl"><span class="gu">### 对比</span></span><span class="cl">加法/减法平方欧氏距离核可以分解，但值域为负，需要额外处理。</span><span class="cl"></span><span class="cl"><span class="gu">### 结论</span></span><span class="cl">选择 $\exp(A) \ast \exp(B)$，并在行方向做归一化。</span><span class="cl">:::</span><span class="cl"></span></code></pre></div></figure>
<p>渲染结果如下：</p>
<section class="cot" data-cot><div class="cot__head"><button class="cot__toggle" type="button" aria-expanded="true"><span class="cot__icon" aria-hidden="true">◈</span><span class="cot__title">为什么选择 Hadamard Exp 核</span><span class="cot__badge">4 步</span><span class="cot__chevron" aria-hidden="true"></span></button><div class="cot__actions"><button type="button" class="cot__action" data-cot-copy>复制</button><button type="button" class="cot__action" data-cot-export>导出</button><button type="button" class="cot__action" data-cot-mindmap>导图</button><button type="button" class="cot__action" data-cot-hide>隐藏</button></div></div><div class="cot__panel"><ol class="cot__steps"><li class="cot__step" data-step="1" data-title="观察"><span class="cot__index" aria-hidden="true">01</span><div class="cot__step-body"><h4 class="cot__step-title">观察</h4><div class="cot__step-content"><p>候选核函数大致可分四类：多项式型、指数型、非负周期型、绝对值型。要满足「精确可分解 + 充分可区分 + 非负 + 几何可解释」四条，可选范围迅速收窄。</p>
</div></div></li><li class="cot__step" data-step="2" data-title="对比"><span class="cot__index" aria-hidden="true">02</span><div class="cot__step-body"><h4 class="cot__step-title">对比</h4><div class="cot__step-content"><p>加法/减法平方欧氏距离核确实可以精确分解为 <span class="math math--inline" data-tex="\|A\|^2 + \|B\|^2 \pm 2A\cdot B">\|A\|^2 + \|B\|^2 \pm 2A\cdot B</span>，但值域含负，作为注意力权重时需要额外处理，且几何解释偏「方向」而弱化「共激活」。</p>
</div></div></li><li class="cot__step" data-step="3" data-title="结论"><span class="cot__index" aria-hidden="true">03</span><div class="cot__step-body"><h4 class="cot__step-title">结论</h4><div class="cot__step-content"><p>选择 <span class="math math--inline" data-tex="\exp(A_i) \ast \exp(B_j)">\exp(A_i) \ast \exp(B_j)</span>。指数变换天然非负、处处可导，并且逐元素乘积刻画的是<strong>特征共激活强度</strong>，与余弦相似度关注的方向一致性互补。</p>
</div></div></li><li class="cot__step" data-step="4" data-title="代价"><span class="cot__index" aria-hidden="true">04</span><div class="cot__step-body"><h4 class="cot__step-title">代价</h4><div class="cot__step-content"><p>指数会放大数值范围，因此需要在特征维度上加约束（例如缩放或减均值），否则长序列上累积量会溢出。</p>
</div></div></li></ol></div></section>
<h3 id="写法二-frontmatter-声明">写法二：Frontmatter 声明</h3>
<p>如果推理步骤不适合插在正文中间，也可以写在 Frontmatter 里：</p>
<figure class="codeblock" data-lang="yaml"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">yaml</span><span class="codeblock__lang">yaml</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-yaml hl"><span class="cl"><span class="nt">cot</span><span class="p">:</span></span><span class="cl"><span class="w">  </span><span class="p p-Indicator">-</span><span class="w"> </span><span class="nt">title</span><span class="p">:</span><span class="w"> </span><span class="l l-Scalar l-Scalar-Plain">为什么要把推理过程单独拎出来</span></span><span class="cl"><span class="w">    </span><span class="nt">body</span><span class="p">:</span><span class="w"> </span><span class="p p-Indicator">|</span></span><span class="cl"><span class="w">      </span><span class="no">正文负责结论，推理过程负责可信度。</span></span><span class="cl"></span></code></pre></div></figure>
<p>这种写法的区块会渲染在正文开头，位置由主题决定。</p>
<p>本文的 Frontmatter 就用了这种写法 —— 你可以在页面顶部看到它。</p>
<h3 id="交互能力">交互能力</h3>
<p>每个 CoT 区块右上角有四个按钮：</p>
<div class="table-wrap"><table class="post-table">
<thead>
<tr>
<th style="text-align:left">按钮</th>
<th style="text-align:left">行为</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align:left">复制</td>
<td style="text-align:left">把标题与全部步骤按 Markdown 格式写入剪贴板</td>
</tr>
<tr>
<td style="text-align:left">导出</td>
<td style="text-align:left">下载为 <code>.md</code> 文件，可直接归档进笔记系统</td>
</tr>
<tr>
<td style="text-align:left">导图</td>
<td style="text-align:left">用纯 SVG 画一张径向思维导图（不依赖任何外部库）</td>
</tr>
<tr>
<td style="text-align:left">隐藏</td>
<td style="text-align:left">临时把这个区块从页面移除，专注读正文</td>
</tr>
</tbody>
</table></div>
<p>点击标题栏可以整体折叠 / 展开。折叠状态下的高度过渡是用 <code>max-height</code> + <code>opacity</code> 做的，比 <code>display: none</code> 更平滑，也比 CSS Grid 的 <code>0fr → 1fr</code> 兼容性更好。</p>
<aside class="admonition admonition--tip"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">✦</span><span class="admonition__title">可选隐藏</span></div><div class="admonition__body">
<p>在容器名后加 <code>[hidden]</code> 可以让区块默认处于折叠状态：</p>
<figure class="codeblock" data-lang="markdown"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">markdown</span><span class="codeblock__lang">markdown</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-markdown hl"><span class="cl">:::cot[hidden] 一段很长的推导</span><span class="cl">...</span><span class="cl"></span></code></pre></div></figure>
</div></aside>
<figure class="codeblock" data-lang="text"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">text</span><span class="codeblock__lang">text</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-text"><span class="cl"> </span><span class="cl"> </span><span class="cl">:::</span><span class="cl"> </span><span class="cl">### 为什么不直接用 `&lt;details&gt;`</span><span class="cl"> </span><span class="cl">原生 `&lt;details&gt;` 能实现折叠，但有三个问题：</span><span class="cl"> </span><span class="cl">1. 没法做高度过渡动画（内容高度未知）；</span><span class="cl">2. 折叠状态的样式在各浏览器上差异较大；</span><span class="cl">3. 想加「导出为 Markdown」这类操作时，得再包一层容器。</span><span class="cl"> </span><span class="cl">自己写一个 `section` 反而更简单 —— 而且反正 CoT 的 HTML 是构建期生成的，没有运行时开销。</span><span class="cl"> </span><span class="cl">### 和正文的关系</span><span class="cl"> </span><span class="cl">CoT 区块在 DOM 上是正文的同级兄弟节点，不在 `&lt;p&gt;` 内部，因此在打印样式里可以整块隐藏：</span><span class="cl"> </span><span class="cl"> </span><span class="cl"> </span><span class="cl">```css</span><span class="cl">@media print {</span><span class="cl">  .cot { display: none; }</span><span class="cl">}</span><span class="cl"> </span></code></pre></div></figure>
<p>这样导出的 PDF 只有结论，不带推理草稿。</p>
<h2 id="几个实现细节">几个实现细节</h2>
<p>渲染管线里有两处不太常见但值得说的做法。</p>
<h3 id="代码块的行号">代码块的行号</h3>
<p>Pygments 的高亮输出里，一个 <code>&lt;span&gt;</code> 可能横跨多行（比如多行字符串）。直接按 <code>\n</code> 切分会切出未闭合的标签，于是需要一层「按行补全标签」的处理：</p>
<figure class="codeblock codeblock--linenos" data-lang="python"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">engine/markdown.py</span><span class="codeblock__lang">python</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-python hl"><span class="cl"><span class="k">def</span><span class="w"> </span><span class="nf">_balance_spans</span><span class="p">(</span><span class="n">html_text</span><span class="p">:</span> <span class="nb">str</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="nb">list</span><span class="p">[</span><span class="nb">str</span><span class="p">]:</span></span><span class="cl"><span class="w">    </span><span class="sd">"""把可能跨行的 Pygments 输出切成行，并在行边界补全未闭合的 &lt;span&gt;。"""</span></span><span class="cl">    <span class="n">lines</span><span class="p">:</span> <span class="nb">list</span><span class="p">[</span><span class="nb">str</span><span class="p">]</span> <span class="o">=</span> <span class="p">[]</span></span><span class="cl">    <span class="n">open_tags</span><span class="p">:</span> <span class="nb">list</span><span class="p">[</span><span class="nb">str</span><span class="p">]</span> <span class="o">=</span> <span class="p">[]</span></span><span class="cl">    <span class="o">...</span></span><span class="cl">    <span class="k">for</span> <span class="n">match</span> <span class="ow">in</span> <span class="n">token_re</span><span class="o">.</span><span class="n">finditer</span><span class="p">(</span><span class="n">html_text</span><span class="p">):</span></span><span class="cl">        <span class="n">chunk</span> <span class="o">=</span> <span class="n">html_text</span><span class="p">[</span><span class="n">pos</span> <span class="p">:</span> <span class="n">match</span><span class="o">.</span><span class="n">start</span><span class="p">()]</span></span><span class="cl">        <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">piece</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">chunk</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="s2">"</span><span class="se">\n</span><span class="s2">"</span><span class="p">)):</span></span><span class="cl">            <span class="k">if</span> <span class="n">i</span><span class="p">:</span></span><span class="cl">                <span class="c1"># 换行处：先把当前行已开启的标签全部闭合</span></span><span class="cl">                <span class="n">lines</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="s2">""</span><span class="o">.</span><span class="n">join</span><span class="p">(</span><span class="n">current</span><span class="p">)</span> <span class="o">+</span> <span class="s2">""</span><span class="o">.</span><span class="n">join</span><span class="p">(</span><span class="s2">"&lt;/span&gt;"</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="n">open_tags</span><span class="p">))</span></span><span class="cl">                <span class="n">current</span> <span class="o">=</span> <span class="p">[</span><span class="n">tag</span> <span class="k">for</span> <span class="n">tag</span> <span class="ow">in</span> <span class="n">open_tags</span><span class="p">]</span></span><span class="cl">            <span class="n">current</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">piece</span><span class="p">)</span></span><span class="cl"></span></code></pre></div></figure>
<h3 id="容器的结束标记">容器的结束标记</h3>
<p>markdown-it 的 <code>paragraph</code> 规则内部会用 <code>state.lineMax</code> 覆盖传入的 <code>endLine</code>。写自定义容器（<code>:::note</code>）时如果不先把 <code>lineMax</code> 收敛到容器结束行，正文就会把收尾的 <code>:::</code> 一起吃进去：</p>
<figure class="codeblock" data-lang="python"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">python</span><span class="codeblock__lang">python</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-python hl"><span class="cl"><span class="n">old_line_max</span> <span class="o">=</span> <span class="n">state</span><span class="o">.</span><span class="n">lineMax</span></span><span class="cl"><span class="n">state</span><span class="o">.</span><span class="n">lineMax</span> <span class="o">=</span> <span class="n">found</span>      <span class="c1"># 收敛到容器结束行，再嵌套解析</span></span><span class="cl"><span class="n">state</span><span class="o">.</span><span class="n">md</span><span class="o">.</span><span class="n">block</span><span class="o">.</span><span class="n">tokenize</span><span class="p">(</span><span class="n">state</span><span class="p">,</span> <span class="n">start_line</span> <span class="o">+</span> <span class="mi">1</span><span class="p">,</span> <span class="n">found</span><span class="p">)</span></span><span class="cl"><span class="n">state</span><span class="o">.</span><span class="n">lineMax</span> <span class="o">=</span> <span class="n">old_line_max</span></span><span class="cl"></span></code></pre></div></figure>
<p>这种坑不写一遍是踩不到的，所以记在这里。</p>
<h2 id="它不做什么">它不做什么</h2>
<p>坦白说清楚边界比罗列功能更有用：</p>
<ul>
<li><strong>不做服务端渲染</strong>：没有 Node、没有 SSR、没有 hydration。归档页、时间线、搜索这些视图是纯客户端从一份 JSON 索引渲染的。</li>
<li><strong>不做数据库</strong>：没有数据库，没有后端。评论走第三方（Giscus），统计走第三方（Umami / Plausible）。</li>
<li><strong>不做富文本编辑器</strong>：写作入口就是 <code>content/posts/&lt;id&gt;/index.md</code>。</li>
</ul>
<h2 id="下一步">下一步</h2>
<ul class="contains-task-list">
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> 主题系统与 <code>ink</code> 默认主题</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> 静态搜索（分片索引 + 拼音 + Worker）</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> 多层次时间线与写作热力图</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox" checked="checked"  type="checkbox"> CoT 折叠展示与思维导图</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox"  type="checkbox"> 图片 AVIF 输出</li>
<li class="task-list-item enabled"><input class="task-list-item-checkbox"  type="checkbox"> 多语言文章互链</li>
</ul>
<p>想看点更实在的，可以读 <a href="../../blogs/exact-linear-attention/app.html" title="wikilink">Exact Linear Attention</a>。</p>
<aside class="admonition admonition--tip"><div class="admonition__head"><span class="admonition__icon" aria-hidden="true">✦</span><span class="admonition__title">想自己跑一遍</span></div><div class="admonition__body">
<figure class="codeblock" data-lang="bash"><figcaption class="codeblock__bar"><span class="codeblock__dot" aria-hidden="true"></span><span class="codeblock__name">bash</span><span class="codeblock__lang">bash</span><button type="button" class="codeblock__copy" data-copy aria-label="复制代码">复制</button></figcaption><div class="codeblock__scroll"><pre><code class="language-bash hl"><span class="cl">pip<span class="w"> </span>install<span class="w"> </span>-r<span class="w"> </span>requirements.txt</span><span class="cl">python<span class="w"> </span>build.py<span class="w"> </span>serve</span><span class="cl"></span></code></pre></div></figure>
<p>构建完成后会自动在 <code>http://127.0.0.1:8000/index.html</code> 打开预览。</p>
</div></aside>
<hr class="footnotes-sep" />
<section class="footnotes">
<ol class="footnotes-list">
<li id="fn1" class="footnote-item"><p>A. Katharopoulos et al., &quot;Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention,&quot; ICML 2020. <a href="#fnref1" class="footnote-backref">↩︎</a></p>
</li>
<li id="fn2" class="footnote-item"><p>I. Schlag, K. Irie, and J. Schmidhuber, &quot;Linear Transformers Are Secretly Fast Weight Programmers,&quot; ICML 2021. <a href="#fnref2" class="footnote-backref">↩︎</a></p>
</li>
</ol>
</section>
]]></content></entry>
<entry><title>安静的机器</title><link href="https://yaunt.github.io/blog/blogs/quiet-machines/app.html"/><id>https://yaunt.github.io/blog/blogs/quiet-machines/app.html</id><updated>2025-11-02T00:00:00Z</updated><published>2025-11-02T00:00:00Z</published><author><name>曦儿</name></author><category term="随笔"/><category term="工程哲学"/><summary type="text">一台安静的机器，通常意味着设计者替使用者做完了大部分决定。 好的构建工具不该让你读完它的文档才能用起来。它应该有一个能跑通的默认值集合，然后在你真的需要时，才把旋钮交出来。 我越来越喜欢这种顺序： 先把…</summary><content type="html"><![CDATA[<p>一台安静的机器，通常意味着设计者替使用者做完了大部分决定。</p>
<p>好的构建工具不该让你读完它的文档才能用起来。它应该有一个能跑通的默认值集合，然后在你真的需要时，才把旋钮交出来。</p>
<p>我越来越喜欢这种顺序：</p>
<ol>
<li>先把最小的闭环跑通；</li>
<li>再把闭环里的每一环做得可控；</li>
<li>最后才考虑把可控的环做成可替换的。</li>
</ol>
<p>反过来做 —— 先设计可替换的抽象，再找地方用 —— 得到的通常是一堆看似灵活、实际没人用的接口。</p>
<blockquote>
<p>复杂性不会消失，它只会被搬到不会伤人的地方。</p>
</blockquote>
<p>这句话我在很多场合引用过。它解释了为什么静态站点生成器比 CMS 更让人安心：把复杂度提前到构建期，运行时就没有东西可以出错了。</p>
<p>当然，代价是构建期要处理的东西变多。这就是增量构建存在的理由 —— 让「提前处理」这件事本身也变得便宜。</p>
<p>至于什么样的复杂度算「不会伤人」，我的判断标准很简单：<strong>出错的时候，能不能在一分钟内定位到具体是哪个文件、哪一行。</strong></p>
]]></content></entry>
<entry><title>开站</title><link href="https://yaunt.github.io/blog/blogs/first-light/app.html"/><id>https://yaunt.github.io/blog/blogs/first-light/app.html</id><updated>2024-03-16T00:00:00Z</updated><published>2024-03-16T00:00:00Z</published><author><name>曦儿</name></author><category term="随笔"/><category term="建站"/><summary type="text">平台换过好几个，文章搬来搬去，最后剩下的只有一堆打不开的链接。 所以这次换个思路：内容就是文件，文件就在仓库里。 写作 = 编辑 发布 = 归档 = GitHub Actions 构建 并推到 没有后台，没有数据库，没有「导出为…</summary><content type="html"><![CDATA[<p>平台换过好几个，文章搬来搬去，最后剩下的只有一堆打不开的链接。</p>
<p>所以这次换个思路：<strong>内容就是文件，文件就在仓库里</strong>。</p>
<ul>
<li>写作 = 编辑 <code>content/posts/&lt;id&gt;/index.md</code></li>
<li>发布 = <code>git push</code></li>
<li>归档 = GitHub Actions 构建 <code>output/</code> 并推到 <code>gh-pages</code></li>
</ul>
<p>没有后台，没有数据库，没有「导出为 Markdown」按钮 —— 因为本来就是 Markdown。</p>
<p>代价是要自己写渲染引擎。但这件事本身就值得写一篇文章，所以也不算亏。</p>
]]></content></entry>
</feed>
