<?xml version="1.0" encoding="UTF-8" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Fathom Layer</title>
    <link>https://fathomlayer.com</link>
    <description>The latest guides, glossaries, and news from Fathom Layer.</description>
    <language>en-us</language>
    <lastBuildDate>Sun, 04 Oct 2026 18:53:45 GMT</lastBuildDate>
    <atom:link href="https://fathomlayer.com/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>What is an LLM inference (serving) engine?</title>
      <link>https://fathomlayer.com/glossary/llm-inference-engine</link>
      <guid>https://fathomlayer.com/glossary/llm-inference-engine</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>An LLM inference engine is the software that loads a language model&apos;s weights, runs the model to generate tokens, and serves the results to applications, usually over an HTTP API. It decides how GPU m...</description>
    </item>
    <item>
      <title>What is a 2nm (N2) process node?</title>
      <link>https://fathomlayer.com/glossary/2nm-process-node</link>
      <guid>https://fathomlayer.com/glossary/2nm-process-node</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>A &quot;2nm&quot; node is a chip-manufacturing generation — not a physical measurement of anything on the chip. It is the industry&apos;s name for a density and efficiency tier. The IEEE&apos;s device-roadmap body (IRDS)...</description>
    </item>
    <item>
      <title>Securing a self-hosted LLM server</title>
      <link>https://fathomlayer.com/guides/securing-a-self-hosted-llm-server</link>
      <guid>https://fathomlayer.com/guides/securing-a-self-hosted-llm-server</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>Running a model locally with Ollama, llama.cpp or vLLM(/glossary/llm-inference-engine) is a few commands. Exposing that server safely is the part the quickstarts skip. The failure mode is consistent:...</description>
    </item>
    <item>
      <title>What is secrets management?</title>
      <link>https://fathomlayer.com/glossary/secrets-management</link>
      <guid>https://fathomlayer.com/glossary/secrets-management</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>Secrets management is the practice of creating, storing, distributing, rotating, revoking and auditing the credentials software needs to run — API keys, database passwords, tokens, SSH keys, certifica...</description>
    </item>
    <item>
      <title>What is prompt injection?</title>
      <link>https://fathomlayer.com/glossary/prompt-injection</link>
      <guid>https://fathomlayer.com/glossary/prompt-injection</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>Prompt injection is an attack where text that reaches a language model is interpreted as an instruction instead of as content — because the model reads instructions and data through the same channel a...</description>
    </item>
    <item>
      <title>vLLM vs SGLang vs Ollama: which serving engine</title>
      <link>https://fathomlayer.com/guides/vllm-vs-sglang-vs-ollama</link>
      <guid>https://fathomlayer.com/guides/vllm-vs-sglang-vs-ollama</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>The question stopped being &quot;which model?&quot; and became &quot;which engine runs it best on my hardware?&quot;. Three names cover most of the decision. Here is how they differ in the way that matters — without benc...</description>
    </item>
    <item>
      <title>What is AI red teaming?</title>
      <link>https://fathomlayer.com/glossary/ai-red-teaming</link>
      <guid>https://fathomlayer.com/glossary/ai-red-teaming</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>AI red teaming is the practice of deliberately attacking your own AI system — with adversarial prompts, jailbreaks and edge-case inputs — to find where it breaks before someone else does. It borrows t...</description>
    </item>
    <item>
      <title>What are FIDO2, WebAuthn and passkeys?</title>
      <link>https://fathomlayer.com/glossary/fido2-webauthn-passkeys</link>
      <guid>https://fathomlayer.com/glossary/fido2-webauthn-passkeys</guid>
      <pubDate>Thu, 03 Sep 2026 19:52:15 GMT</pubDate>
      <description>FIDO2 is a set of standards for logging in with public-key cryptography instead of a password; WebAuthn is the browser API half of it; a passkey is the credential it creates. Together they are what &quot;p...</description>
    </item>
    <item>
      <title>What is MCP tool poisoning?</title>
      <link>https://fathomlayer.com/glossary/mcp-tool-poisoning</link>
      <guid>https://fathomlayer.com/glossary/mcp-tool-poisoning</guid>
      <pubDate>Thu, 03 Sep 2026 18:16:55 GMT</pubDate>
      <description>MCP tool poisoning is an attack where a malicious or compromised MCP(/glossary/mcp) server hides instructions in the metadata a model reads — tool descriptions, parameter schemas or tool output — inst...</description>
    </item>
    <item>
      <title>What is the A2A (Agent2Agent) Protocol?</title>
      <link>https://fathomlayer.com/glossary/a2a-agent2agent-protocol</link>
      <guid>https://fathomlayer.com/glossary/a2a-agent2agent-protocol</guid>
      <pubDate>Thu, 03 Sep 2026 18:16:55 GMT</pubDate>
      <description>The Agent2Agent (A2A) Protocol is an open standard for letting independent AI agents — built by different teams, vendors or frameworks — discover each other and delegate work, without a bespoke integr...</description>
    </item>
    <item>
      <title>What are MLA and DeepSeekMoE?</title>
      <link>https://fathomlayer.com/glossary/deepseek-mla-and-deepseekmoe</link>
      <guid>https://fathomlayer.com/glossary/deepseek-mla-and-deepseekmoe</guid>
      <pubDate>Thu, 03 Sep 2026 18:16:55 GMT</pubDate>
      <description>Multi-head Latent Attention (MLA) and DeepSeekMoE are the two architecture choices that let DeepSeek-V3 run as a 671-billion-parameter model while activating only about 37 billion parameters per token...</description>
    </item>
    <item>
      <title>What are ternary LLMs (BitNet b1.58)?</title>
      <link>https://fathomlayer.com/glossary/ternary-llms-bitnet-b1-58</link>
      <guid>https://fathomlayer.com/glossary/ternary-llms-bitnet-b1-58</guid>
      <pubDate>Thu, 03 Sep 2026 18:16:55 GMT</pubDate>
      <description>A ternary LLM constrains every weight to one of three values — −1, 0 or +1 — which needs about 1.58 bits of storage each (log₂3 ≈ 1.58), against 16 bits for a standard model. The idea comes from Micro...</description>
    </item>
    <item>
      <title>Optimising local LLM inference: where the bottleneck actually is</title>
      <link>https://fathomlayer.com/guides/local-llm-hardware-optimization-guide</link>
      <guid>https://fathomlayer.com/guides/local-llm-hardware-optimization-guide</guid>
      <pubDate>Sat, 29 Aug 2026 18:47:13 GMT</pubDate>
      <description>Local inference is memory-bandwidth bound, not compute bound. Generating one token requires streaming every active weight from memory exactly once. So the ceiling is:  text max tokens/sec ≈ memoryband...</description>
    </item>
    <item>
      <title>HNSW (Hierarchical Navigable Small World)</title>
      <link>https://fathomlayer.com/glossary/hnsw-algorithm</link>
      <guid>https://fathomlayer.com/glossary/hnsw-algorithm</guid>
      <pubDate>Sat, 29 Aug 2026 18:34:00 GMT</pubDate>
      <description>HNSW (Hierarchical Navigable Small World) is a graph index for approximate nearest-neighbour search: it finds vectors close to a query by walking a layered graph instead of comparing the query with ev...</description>
    </item>
    <item>
      <title>Quantization (LLM)</title>
      <link>https://fathomlayer.com/glossary/quantization</link>
      <guid>https://fathomlayer.com/glossary/quantization</guid>
      <pubDate>Sat, 29 Aug 2026 18:34:00 GMT</pubDate>
      <description>Quantization stores a model&apos;s weights at lower numeric precision — 8, 4, or even 2 bits instead of 16 — trading a small, measurable quality loss for a large cut in memory and bandwidth. It is the tech...</description>
    </item>
    <item>
      <title>KV Cache (Key-Value Cache)</title>
      <link>https://fathomlayer.com/glossary/kv-cache</link>
      <guid>https://fathomlayer.com/glossary/kv-cache</guid>
      <pubDate>Sat, 29 Aug 2026 18:33:58 GMT</pubDate>
      <description>The KV cache stores the key and value vectors the attention mechanism has already computed for every token in the context, so the model never recomputes them. It is what makes generation after the fir...</description>
    </item>
    <item>
      <title>What is a Local AI Appliance?</title>
      <link>https://fathomlayer.com/glossary/local-ai-appliance</link>
      <guid>https://fathomlayer.com/glossary/local-ai-appliance</guid>
      <pubDate>Mon, 17 Aug 2026 19:40:01 GMT</pubDate>
      <description>A local AI appliance is a self-contained computer bought mainly to run AI models on your own desk or network, rather than renting them from a cloud API. It is a descriptive label, not a standard indus...</description>
    </item>
    <item>
      <title>What is MicroLED, and why does it matter for wearables?</title>
      <link>https://fathomlayer.com/glossary/microled-wearables</link>
      <guid>https://fathomlayer.com/glossary/microled-wearables</guid>
      <pubDate>Mon, 17 Aug 2026 19:40:01 GMT</pubDate>
      <description>MicroLED is a display technology in which every pixel, or every subpixel, is its own microscopic inorganic LED, so the panel makes its own light without a backlight. In wearables it matters in two dif...</description>
    </item>
    <item>
      <title>Llama 3 vs DeepSeek locally: which open model to run</title>
      <link>https://fathomlayer.com/guides/llama-3-vs-deepseek-local</link>
      <guid>https://fathomlayer.com/guides/llama-3-vs-deepseek-local</guid>
      <pubDate>Mon, 17 Aug 2026 19:40:01 GMT</pubDate>
      <description>Both are strong open-weight families you can run locally. The split is simple: DeepSeek&apos;s reasoning models think before answering (a visible chain-of-thought pass that costs tokens and latency), while...</description>
    </item>
    <item>
      <title>Building a Tesla P40 rig for local AI: what you&apos;re really signing up for</title>
      <link>https://fathomlayer.com/guides/tesla-p40-local-ai-rig</link>
      <guid>https://fathomlayer.com/guides/tesla-p40-local-ai-rig</guid>
      <pubDate>Mon, 17 Aug 2026 19:40:00 GMT</pubDate>
      <description>The Tesla P40 is the cheapest route to a large VRAM pool: 24 GB per card on the used market for a fraction of a consumer 24 GB GPU. Two of them give you 48 GB for the price of one modern card. The cat...</description>
    </item>
  </channel>
</rss>