<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Usman Shahid - Writing</title><description>Technical writeups and opinions on local LLMs, agent orchestration, and small-model agents.</description><link>https://codemug.github.io/</link><item><title>An SRE&apos;s guide to deploying Large Language Models, Part 2: from a base model to a served one</title><link>https://codemug.github.io/blog/sre-guide-deploying-llms-part-2/</link><guid isPermaLink="true">https://codemug.github.io/blog/sre-guide-deploying-llms-part-2/</guid><description>How a raw next-token predictor is turned into a model that follows instructions, reasons, and stays in bounds, and the GPU, memory, and quantization mechanics that decide whether serving it is fast or ruinously slow.</description><pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate><category>llms</category><category>sre</category><category>fine-tuning</category><category>rlhf</category><category>inference</category><category>quantization</category></item><item><title>An SRE&apos;s guide to deploying Large Language Models, Part 1: understand the workload</title><link>https://codemug.github.io/blog/sre-guide-deploying-llms-part-1/</link><guid isPermaLink="true">https://codemug.github.io/blog/sre-guide-deploying-llms-part-1/</guid><description>Before you can deploy and serve LLMs reliably, you have to understand them as a workload. A ground-up tour of the transformer, from tokens to attention to the feed-forward step, for SREs.</description><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate><category>llms</category><category>sre</category><category>transformers</category><category>inference</category></item><item><title>I created Primer, a context optimized, loop engineering AI platform so I can get meaningful work done on my 16GB VRAM</title><link>https://codemug.github.io/blog/gaming-pc-agent-platform/</link><guid isPermaLink="true">https://codemug.github.io/blog/gaming-pc-agent-platform/</guid><description>I bought a 16GB gaming GPU, tried to get real work out of small local models, and ended up building Primer: a platform for running lots of tiny, focused agents instead of prompting one big one.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate><category>local-llms</category><category>agents</category><category>primer</category><category>context-engineering</category></item></channel></rss>