Skip to content

Archive

Transformer

4 articles
Tech 19 Sep 2026 5 min read

When a Transformer Surface Can Become an Electric-Shock Hazard

A transformer can have exposed metal around its core, frame, mounting hardware, or enclosure while carrying hazardous voltage on its windings. Those metal parts are not automatically live simply because the transformer is energized. Whether touching them can produce an electric shock depends on the transformer’s construction, insulation condition, protective earthing, and the electrical path between the person and the circuit. The important distinction is between a conductor intended to carry voltage and an accessible conductive part that should remain separated from it.

Artificial Intelligence 03 Sep 2026 6 min read

Self-Attention in Transformer Models

Transformers can process relationships between tokens without stepping through a sequence one token at a time. The mechanism that makes this possible is self-attention: each token builds a weighted view of other tokens in the same context. The formula is compact. The sections below show what the calculation does, how masking sets the information boundary, and where the computational cost comes from. Start with token representations Before attention runs, each input token is represented by a vector. Let the matrix X contain those token representations. A transformer layer applies learned projections to produce three matrices:

Artificial Intelligence 03 Sep 2026 10 min read

Positional Information in Transformer Models

Self-attention can compare every token with other tokens in a context, but the comparison alone does not tell the model where those tokens occur. A sentence is not just a collection of words: changing their order can change the meaning. Transformer models therefore need a way to represent positional information. This mechanism lets the network distinguish, for example, the first occurrence of a token from a later occurrence and reason about relationships such as “the previous token” or “far earlier in the document.”

Artificial Intelligence 03 Sep 2026 8 min read

Activation Functions in Transformer Feed-Forward Networks

Attention gets much of the attention in transformer explanations, but every transformer layer also contains a feed-forward network that performs substantial computation on each token representation. The activation function inside that network is a small-looking design choice with an important job: it introduces nonlinearity so the network can learn transformations that stacked linear projections alone cannot express. The feed-forward block matters when reading model architectures, comparing implementations, estimating parameter and compute costs, or deciding whether two designs are actually equivalent.