Skip to content

Instantly share code, notes, and snippets.

View magnus919's full-sized avatar
💭
Diving back into Hugo and figuring out how/where I want to deploy it.

Magnus Hedemark magnus919

💭
Diving back into Hugo and figuring out how/where I want to deploy it.
View GitHub Profile
@magnus919
magnus919 / bytedance-ft-de-spin.md
Created August 8, 2026 01:26
Evidence audit of the Financial Times report on ByteDance's possible 10T model

De-spinning the Financial Times report on ByteDance's possible 10T-parameter model

Review date: 2026-08-07 EDT (2026-08-08 UTC)

Source and decision boundary

The message is trying to move the reader toward this belief: ByteDance is building a model so large that China is about to match or surpass the leading US labs, and model scale is evidence of that coming capability.

The supplied copy was a FreeDium mirror of a Financial Times article. The canonical FT page exists and is titled “ByteDance targets mega AI model nearing Anthropic’s Mythos,” but it required a subscription at access time.[1] This audit uses the mirrored article text only to identify what FT reported, then checks material claims against direct company, model, data-provider, and independent reporting sources. It does not reproduce the article.

@magnus919
magnus919 / forward-deployed-engineering-demo.md
Last active August 6, 2026 19:56
Forward-deployed engineering, demonstrated: what a 3-epoch SkillOpt-optimized agent skill does — vague sponsor, mid-stream join, AI release gate, routing — with real validation output

Forward-Deployed Engineering, Demonstrated

An agent skill that carries an embedded technical engagement from a vague stakeholder need to an adopted, measured outcome — and stops the engagement when the evidence or authority isn't there.

Get the skill: forward-deployed-engineering in the agent-skills repository (SKILL.md, manifest, 8 references, 10 templates, 15 eval cases; optimized in PR #294).

@magnus919
magnus919 / README.md
Created August 5, 2026 17:35
dsm5 Agent Skill demo: what an AI says when someone asks about mental health (real conversations, real captures)

dsm5: what an AI says when someone asks about mental health

What this is: dsm5 is an Agent Skill that gives an AI a disciplined way to talk about mental health. It is a paraphrased companion to the DSM-5-TR (American Psychiatric Association, 2022): a reference library of diagnostic criteria, specifiers, and differentials, wrapped in rules that keep the conversation safe and honest.

Why it matters: an AI without guardrails will happily answer "do I have bipolar disorder?" with a confident verdict. That is dangerous. This skill makes the AI do six things on every answer: triage safety first, use calibrated language ("consistent with" rather than "you have"), compare the presentation against actual criteria with met/unmet/unknown kept separate, reason through the differential, give concrete next steps, and end by handing the decision to a qualified clinician.

Everything below is a real conversation between a person and an AI running t

@magnus919
magnus919 / agent-skills-v0.6.0-release-announcement.md
Created August 3, 2026 02:09
agent-skills v0.6.0 release announcement: 20 new skill entry points, binary analysis, product and production operations, QA, and release engineering

agent-skills v0.6.0: a bigger, safer toolkit for building and operating AI agents

agent-skills v0.6.0 is out.

This is the release where the collection starts to feel less like a directory of prompts and more like an operating toolkit for serious agent work: reusable methods, executable CLIs, evidence contracts, quality gates, lifecycle handoffs, and production controls.

The release at a glance

  • 20 new skill entry points, including three orchestration bundles
  • 8 existing skill surfaces updated
@magnus919
magnus919 / nous-portal-review.md
Created August 2, 2026 06:42
Nous Portal public site UX, accessibility, and product review (2026-08-02)

Nous Portal Public Site Review

Review date: 2026-08-02 06:39 UTC Reviewer: Independent exploratory UX, accessibility, and product review Target: https://portal.nousresearch.com/

Executive assessment

Nous Portal presents a distinctive, technically credible front door to Hermes Agent. The public site communicates a coherent bundle: model access, hosted tools, Hermes Cloud, API access, and a shared credit balance. The strongest parts are the authored visual identity, the breadth of the public model catalog, the working search and plan controls, the useful API reference, and the cloud FAQ.

@magnus919
magnus919 / gpt56-family-bang-for-buck.md
Created August 2, 2026 03:01
GPT-5.6 Sol, Terra, Luna, and DeepSeek V4 Flash: benchmarked cost-effectiveness

GPT-5.6 Sol, Terra, and Luna: which one buys the most completed work?

Research date: 2026-08-01

Executive answer

The strongest currently available public evidence does not support using Terra as the default solely for cost-effectiveness. Artificial Analysis, which evaluated all three GPT-5.6 tiers across reasoning efforts, reports that Luna and Sol are always on its intelligence-versus-cost Pareto frontier ahead of Terra. In its words: for any Terra effort level, a Luna or Sol configuration is either more intelligent at no added cost or equally intelligent at lower cost.

That is not a blanket instruction to replace Terra with Luna-max. The main trade-off is latency: in a direct current comparison, Luna-max scores 51 on Artificial Analysis’s Intelligence Index versus 46 for Terra-medium, while costing less and generating tokens faster, but its time to first token is 126.28 seconds versus Terra-medium’s 1.60 seconds. For an interactive task where that first response matters, Ter

@magnus919
magnus919 / gpt-5-6-luna-max-vs-sol-medium.md
Last active August 2, 2026 02:48
GPT-5.6 Luna max vs. Sol medium: full source-checked assessment

GPT-5.6 Luna (max) vs. Sol (medium): a source-checked comparison

Bottom line

False. The claim “Luna at highest reasoning is smarter than Sol Medium” is contradicted by the directly relevant configuration comparison from Artificial Analysis.

On its Artificial Analysis Intelligence Index, where higher is better:

Configuration Score
@magnus919
magnus919 / README.md
Created August 2, 2026 01:36
QA Methodology 2.0: feature brief to release-ready QA plan

A QA skill should change the plan before code ships

This is a compact, runnable case study for QA Methodology 2.0, an MIT-licensed, platform-agnostic Agent Skill for QA leads, SDETs, developers, and agentic software teams.

The useful question is not “can it calculate a risk score?” It is: can it turn an ambiguous feature request into an auditable plan for deciding whether that feature is safe to release?

This demo starts with a deliberately generic request for a workspace-data export API. It then applies the skill to produce five concrete QA artifacts:

Starting point QA Methodology output Why it matters
@magnus919
magnus919 / README.md
Last active July 31, 2026 22:20
A live static-analysis skill demo for AI agents

A static-analysis skill an AI agent can actually interrogate

What this is: binary-analysis is an Agent Skill that gives an AI agent a deterministic, read-only workflow for inspecting unfamiliar PE, ELF, and Mach-O files using Ghidra-backed static analysis. It can identify binary structure, enumerate imports and strings, decompile functions, explore references/call paths, run bounded heuristics, and export evidence-backed reports.

Get it from: the magnus919/agent-skills GitHub repository. The skill lives at binary-analysis/.

An agent handling an unfamiliar executable needs more than strings, a one-line verdict, or a polished hallucination. It needs a bounded, reproducible workflow that can preserve evidence, name uncertainty, and leave behind an audit trail.

This demo uses a harmless, transparent Mach-O f

@magnus919
magnus919 / fact-check-ivermectin-post.md
Created July 29, 2026 23:34
Fact-check: ivermectin, Fauci, and COVID-19 vaccines

Fact-check: ivermectin, Fauci, and COVID-19 vaccines

Verdict: Misleading.

The post combines one substantively supportable observation, that Anthony Fauci publicly discouraged ivermectin for COVID-19, with an unsupported accusation about his motive, untraceable headline statistics, and a false classification of mRNA vaccines as “gene therapies.” It presents the resulting conspiracy narrative as though it follows from the evidence. It does not.

What was checked

The screenshot is a post by Nicolas Hulscher, MPH (@NicHulscher). It says: