Skip to content

Instantly share code, notes, and snippets.

@simonw
Created August 16, 2026 18:45
Show Gist options
  • Select an option

  • Save simonw/491e55ac9d741202ea0af5d9d93775d4 to your computer and use it in GitHub Desktop.

Select an option

Save simonw/491e55ac9d741202ea0af5d9d93775d4 to your computer and use it in GitHub Desktop.

Session Transcript

  • Session ID: 01a00be0-8f53-7edc-90fa-7aabc25fb36a
  • Model: spark/qwen3.8-27b
  • Working dir: /private/tmp/transform
  • Thinking level: medium

πŸ’¬ User

Write Python code to convert this jsonl to markdown

🧠 Thinking

The user wants me to write Python code that converts JSONL to Markdown. First, let's check the current working directory and see if there are any JSONL files available as examples.

πŸ”§ bash

{
  "command": "ls -la && find . -maxdepth 2 -name \"*.jsonl\" 2>/dev/null | head"
}

↩️ bash result

total 112
drwxr-xr-x    3 simon  wheel     96 Aug 16 11:40 .
drwxrwxrwt  161 root   wheel   5152 Aug 16 11:40 ..
-rw-r--r--    1 simon  wheel  56745 Aug 16 11:40 2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl
./2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl
🧠 Thinking

There's a JSONL file. Let me look at its structure and understand what kind of data it contains β€” it looks like a session log (based on the filename pattern, probably from an agent harness).

πŸ”§ bash

{
  "command": "head -c 3000 \"2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl\" && echo \"---LINECOUNT---\" && wc -l \"2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl\""
}

↩️ bash result

{"type":"session","version":3,"id":"01a006d4-fa13-76b4-83f3-1a0ab01392c5","timestamp":"2026-08-15T19:10:16.083Z","cwd":"/Users/simon/Dropbox/dev/datasette"}
{"type":"model_change","id":"5c7ae43b","parentId":null,"timestamp":"2026-08-15T19:10:16.115Z","provider":"spark","modelId":"qwen3.8-27b"}
{"type":"thinking_level_change","id":"46634c4a","parentId":"5c7ae43b","timestamp":"2026-08-15T19:10:16.115Z","thinkingLevel":"low"}
{"type":"message","id":"cf42aa0d","parentId":"46634c4a","timestamp":"2026-08-15T19:10:18.619Z","message":{"role":"user","content":[{"type":"text","text":"how does auth work?"}],"timestamp":1786821018618}}
{"type":"message","id":"1aca50ca","parentId":"cf42aa0d","timestamp":"2026-08-15T19:10:30.414Z","message":{"role":"assistant","content":[{"type":"thinking","thinking":"The user is asking \"How does authentication work?\" in the context of the Datasette project. Let's look at the codebase and understand how authentication works in Datasette.\n\nFirst, let's start by exploring the repository structure.\n","thinkingSignature":"{\"id\":\"rs_lnnqaighkvjjwmc7o5f8a\",\"type\":\"reasoning\",\"status\":\"completed\",\"summary\":[],\"content\":[{\"type\":\"reasoning_text\",\"text\":\"The user is asking \\\"How does authentication work?\\\" in the context of the Datasette project. Let's look at the codebase and understand how authentication works in Datasette.\\n\\nFirst, let's start by exploring the repository structure.\\n\"}]}"},{"type":"toolCall","id":"call_3892970073740005|fc_jyg2ancmw2cd11e2sltbt","name":"bash","arguments":{"command":"ls /Users/simon/Dropbox/dev/datasette && ls /Users/simon/Dropbox/dev/datasette/datasette 2>/dev/null | head -50"}}],"api":"openai-responses","provider":"spark","model":"qwen3.8-27b","usage":{"input":1606,"output":48,"cacheRead":0,"cacheWrite":0,"reasoning":48,"totalTokens":1654,"cost":{"input":0,"output":0,"cacheRead":0,"cacheWrite":0,"total":0}},"stopReason":"toolUse","timestamp":1786821018637,"responseId":"resp_84ecfc7657df80970c4e2709ea6beea35be7b0f8ca269a84","rawStopReason":"completed"}}
{"type":"message","id":"cd4672ed","parentId":"1aca50ca","timestamp":"2026-08-15T19:10:30.443Z","message":{"role":"toolResult","toolCallId":"call_3892970073740005|fc_jyg2ancmw2cd11e2sltbt","toolName":"bash","content":[{"type":"text","text":"1.0-release-notes.md\n1.0-release-report.md\nbenchmarks\nblogs.db\nbuild\nCODE_OF_CONDUCT.md\ncodecov.yml\ncontent.db\ndata.db\ndatasette\ndatasette.egg-info\ndemos\ndist\nDockerfile\ndocs\nfixtures.db\nhello.db\nhtmlcov\nignored\ninternal-edit-demo.db\ninternal-operation-sql.db\ninternal-queries.db\ninternal-sql-write.db\ninternal.db\nissues.db\njump-sections-demo-plugin\nJustfile\nLICENSE\nMANIFEST.in\nnode_modules\npackage-lock.json\npackage.json\nPipfile\npyproject.toml\npytest.ini\npython-3.14-venv\npython3.14-venv\nREADME.md\nruff.toml\nscratch\nsetup.cfg\ntest-in-pyodide-with-shot-scraper.sh\ntests\ntmp-vacuum-copy2.db\ntmp.db\nuv.lock\nvenv\n__init__.py\n__main__.py\n__py---LINECOUNT---
      25 2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl
🧠 Thinking

This is a pi session log JSONL file. Let me look at more lines to understand the full structure β€” especially what content types exist (text, thinking, toolCall, toolResult).

πŸ”§ bash

{
  "command": "python3 - <<'EOF'\nimport json\nwith open(\"2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl\") as f:\n    for i, line in enumerate(f):\n        d = json.loads(line)\n        t = d.get(\"type\")\n        if t == \"message\":\n            m = d[\"message\"]\n            role = m.get(\"role\")\n            kinds = [c.get(\"type\") for c in m.get(\"content\", [])] if isinstance(m.get(\"content\"), list) else type(m.get(\"content\")).__name__\n            tools = [c.get(\"name\") for c in (m.get(\"content\") or []) if isinstance(c, dict) and c.get(\"type\") == \"toolCall\"] if isinstance(m.get(\"content\"), list) else []\n            print(i, t, role, kinds, tools)\n        else:\n            print(i, t, {k: d[k] for k in d if k not in (\"id\",\"parentId\",\"timestamp\")})\nEOF"
}

↩️ bash result

0 session {'type': 'session', 'version': 3, 'cwd': '/Users/simon/Dropbox/dev/datasette'}
1 model_change {'type': 'model_change', 'provider': 'spark', 'modelId': 'qwen3.8-27b'}
2 thinking_level_change {'type': 'thinking_level_change', 'thinkingLevel': 'low'}
3 message user ['text'] []
4 message assistant ['thinking', 'toolCall'] ['bash']
5 message toolResult ['text'] []
6 message assistant ['thinking', 'toolCall', 'toolCall'] ['read', 'bash']
7 message toolResult ['text'] []
8 message toolResult ['text'] []
9 message assistant ['thinking', 'toolCall'] ['bash']
10 message toolResult ['text'] []
11 message assistant ['thinking', 'toolCall', 'toolCall'] ['read', 'bash']
12 message toolResult ['text'] []
13 message toolResult ['text'] []
14 message assistant ['thinking', 'toolCall'] ['bash']
15 message toolResult ['text'] []
16 message assistant ['thinking', 'toolCall'] ['bash']
17 message toolResult ['text'] []
18 message assistant ['thinking', 'toolCall'] ['bash']
19 message toolResult ['text'] []
20 message assistant ['thinking', 'toolCall'] ['bash']
21 message toolResult ['text'] []
22 message assistant ['thinking', 'toolCall'] ['bash']
23 message toolResult ['text'] []
24 message assistant ['thinking', 'text'] []
🧠 Thinking

The structure is clear: this is a pi agent session log containing the following:

  • session metadata (id, cwd)
  • model_change, thinking_level_change events
  • message entries with role being user/assistant/toolResult
    • Content blocks: text, thinking, toolCall

I'll write a Python script that converts this to clean markdown. The output should include the following:

  1. A header containing session info (id, cwd, model)
  2. User messages as blockquotes or sections
  3. Assistant text as normal paragraphs
  4. Thinking blocks β€” optional, probably collapsible <details> elements
  5. Tool calls in fenced code blocks with language specified
  6. Tool results in fenced blocks

Let me write a clean, reusable script jsonl_to_md.py and run it against this file to verify the output looks good.

This is a pi agent session log (user/assistant/tool messages with thinking + tool calls). I'll write a converter and test it on your file.

πŸ”§ write

{
  "path": "/private/tmp/transform/jsonl_to_md.py",
  "content": "#!/usr/bin/env python3\n\"\"\"Convert a pi agent session JSONL log to readable Markdown.\n\nUsage:\n    python jsonl_to_md.py input.jsonl [-o output.md] [--no-thinking]\n\nHandles line types found in pi session logs:\n  - session / model_change / thinking_level_change (metadata header)\n  - message lines with roles: user, assistant, toolResult\n    content blocks: text, thinking, toolCall\n\"\"\"\n\nimport argparse\nimport json\nimport re\nfrom pathlib import Path\n\n\ndef fence(text: str, lang: str = \"\") -> str:\n    \"\"\"Wrap text in a fenced code block, widening the fence if needed.\"\"\"\n    max_fence = max([len(m) for m in re.findall(r\"^`{3,}\", text, flags=re.M)] + [0])\n    marker = \"`\" * max(4, max_fence + 1)\n    return f\"{marker}{lang}\\n{text.rstrip()}\\n{marker}\"\n\n\ndef render_tool_call(block: dict) -> str:\n    name = block.get(\"name\", \"tool\")\n    args = block.get(\"arguments\", {})\n    body = json.dumps(args, indent=2, ensure_ascii=False) if isinstance(args, (dict, list)) else str(args)\n    lang = {\"bash\": \"bash\", \"read\": \"\", \"write\": \"\", \"edit\": \"\"}.get(name, \"\")\n    return f\"**πŸ”§ `{name}`**\\n\\n{fence(body, lang)}\"\n\n\ndef render_tool_result(block: dict) -> str:\n    name = block.get(\"toolName\") or \"result\"\n    content = block.get(\"content\", [])\n    if isinstance(content, list):\n        text = \"\\n\".join(c.get(\"text\", \"\") for c in content if isinstance(c, dict))\n    else:\n        text = str(content)\n    isError = block.get(\"isError\") or (block.get(\"content\") and isinstance(block[\"content\"], list)\n                                      and any(isinstance(c, dict) and c.get(\"type\") == \"error\" for c in block[\"content\"]))\n    label = f\"**❌ {name} failed**\" if isError else f\"**↩️ `{name}` result**\"\n    return f\"{label}\\n\\n{fence(text or '(empty)', '')}\"\n\n\ndef render_message(d: dict, include_thinking: bool) -> list[str]:\n    \"\"\"Render one 'message' line into a list of markdown chunks.\"\"\"\n    m = d[\"message\"]\n    role = m.get(\"role\")\n    content = m.get(\"content\", [])\n    if not isinstance(content, list):\n        content = [{\"type\": \"text\", \"text\": str(content)}]\n\n    parts: list[str] = []\n    if role == \"user\":\n        text = \"\\n\\n\".join(c.get(\"text\", \"\") for c in content if c.get(\"type\") == \"text\").strip()\n        parts.append(f\"## πŸ’¬ User\\n\\n{text}\")\n    elif role == \"assistant\":\n        body: list[str] = []\n        for c in content:\n            ctype = c.get(\"type\")\n            if ctype == \"thinking\" and include_thinking:\n                body.append(f\"<details>\\n<summary>🧠 Thinking</summary>\\n\\n{c.get('thinking', '').strip()}\\n\\n</details>\")\n            elif ctype == \"text\":\n                body.append(c.get(\"text\", \"\").strip())\n            elif ctype == \"toolCall\":\n                body.append(render_tool_call(c))\n        if body:\n            parts.append(\"\\n\\n\".join(body))\n    elif role == \"toolResult\":\n        for c in content if isinstance(content, list) else [content]:\n            # toolResult lines carry their own fields; handle both shapes\n            src = c if (c.get(\"toolCallId\") or c.get(\"toolName\")) else m\n            parts.append(render_tool_result(src))\n    return parts\n\n\ndef convert(path: Path, include_thinking: bool) -> str:\n    out: list[str] = []\n    session_meta: dict = {}\n\n    with path.open(encoding=\"utf-8\") as f:\n        for line in f:\n            line = line.strip()\n            if not line:\n                continue\n            try:\n                d = json.loads(line)\n            except json.JSONDecodeError:\n                out.append(f\"> ⚠️ Skipped invalid JSONL line: `{line[:80]}…`\\n\")\n                continue\n\n            t = d.get(\"type\")\n            if t == \"session\":\n                session_meta.update({k: d[k] for k in (\"id\", \"cwd\", \"version\") if k in d})\n            elif t == \"model_change\" and \"modelId\" in d:\n                session_meta[\"model\"] = f\"{d.get('provider', '')}/{d['modelId']}\" if d.get(\"provider\") else d[\"modelId\"]\n            elif t == \"thinking_level_change\":\n                session_meta.setdefault(\"thinkingLevel\", d.get(\"thinkingLevel\"))\n            elif t == \"message\":\n                out.extend(render_message(d, include_thinking))\n\n    # Header\n    header = [\"# Session Transcript\", \"\"]\n    for label, key in [(\"Session ID\", \"id\"), (\"Model\", \"model\"),\n                       (\"Working dir\", \"cwd\"), (\"Thinking level\", \"thinkingLevel\")]:\n        if session_meta.get(key):\n            v = str(session_meta[key])\n            header.append(f\"- **{label}:** `{v}`\" if key != \"id\" else f\"- **{label}:** {v}\")\n    out_md = \"\\n\\n\".join(out)\n    return \"\\n\".join(header).rstrip() + \"\\n\\n---\\n\\n\" + (out_md.rstrip() or \"*No messages found.*\")\n\n\ndef main() -> None:\n    ap = argparse.ArgumentParser(description=__doc__)\n    ap.add_argument(\"input\", type=Path, help=\"Input .jsonl session file\")\n    ap.add_argument(\"-o\", \"--output\", type=Path, default=None,\n                    help=\"Output .md path (default: <input>.md)\")\n    ap.add_argument(\"--no-thinking\", action=\"store_true\",\n                    help=\"Omit assistant thinking blocks\")\n    args = ap.parse_args()\n\n    out_path = args.output or args.input.with_suffix(\".md\")\n    md = convert(args.input, include_thinking=not args.no_thinking)\n    out_path.write_text(md + \"\\n\", encoding=\"utf-8\")\n    print(f\"Wrote {out_path} ({len(md)} chars)\")\n\n\nif __name__ == \"__main__\":\n    main()\n"
}

↩️ write result

Successfully wrote 5321 bytes to /private/tmp/transform/jsonl_to_md.py
🧠 Thinking

Wait, there's a problem with the rendering of toolResult. Looking at the JSONL structure:

{"type":"message","id":"cd4672ed",...,"message":{"role":"toolResult","toolCallId":"...","toolName":"bash","content":[{"type":"text","text":"..."}]}}

So m itself is a dict with role=toolResult, toolCallId, toolName, and content. My render_message for the toolResult case: for c in content if isinstance(content, list) else [content]: src = c if (c.get("toolCallId") or c.get("toolName")) else m β€” this works because the text blocks inside content don't have toolCallId/toolName, so src falls back to m. OK, no problem.

Also, I need to check whether isError is used in pi's session format β€” it might be "isError" on the message or on the result block. Let me try running the script and checking the output.

Now let's run it against your file and inspect the result:

πŸ”§ bash

{
  "command": "cd /private/tmp/transform && python3 jsonl_to_md.py \"2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl\" -o transcript.md && head -60 transcript.md"
}

↩️ bash result

Wrote transcript.md (39514 chars)
# Session Transcript

- **Session ID:** 01a006d4-fa13-76b4-83f3-1a0ab01392c5
- **Model:** `spark/qwen3.8-27b`
- **Working dir:** `/Users/simon/Dropbox/dev/datasette`
- **Thinking level:** `low`

---

## πŸ’¬ User

how does auth work?

<details>
<summary>🧠 Thinking</summary>

The user is asking "How does authentication work?" in the context of the Datasette project. Let's look at the codebase and understand how authentication works in Datasette.

First, let's start by exploring the repository structure.

</details>

**πŸ”§ `bash`**

````bash
{
  "command": "ls /Users/simon/Dropbox/dev/datasette && ls /Users/simon/Dropbox/dev/datasette/datasette 2>/dev/null | head -50"
}
````

**↩️ `bash` result**

````
1.0-release-notes.md
1.0-release-report.md
benchmarks
blogs.db
build
CODE_OF_CONDUCT.md
codecov.yml
content.db
data.db
datasette
datasette.egg-info
demos
dist
Dockerfile
docs
fixtures.db
hello.db
htmlcov
ignored
internal-edit-demo.db
internal-operation-sql.db
internal-queries.db
internal-sql-write.db
internal.db
issues.db
jump-sections-demo-plugin
Justfile
🧠 Thinking

It works. Let me also check the end of the file to make sure the final assistant message is being rendered correctly, and verify that there are no issues with the toolResult isError handling (the pi format uses isError at the message level β€” my code checks block.get("isError"), but here block = m, which contains role/toolCallId/toolName/content. Let me check whether any line in this file has isError).

πŸ”§ bash

{
  "command": "cd /private/tmp/transform && tail -25 transcript.md && echo \"=== isError present? ===\" && grep -c 'isError' *.jsonl || true"
}

↩️ bash result

If a token handler raises `TokenInvalid`, that's an immediate 401 even if other credentials would have worked. Unauthenticated requests get `actor=None`.

## 2. Cookie auth β€” the `ds_actor` cookie (`datasette/actor_auth_cookie.py`)

The core ships one built-in credential: a signed session cookie.

- **Setting**: plugins call `set_actor_cookie(response, actor)` (app.py ~2451), which does `response.set_cookie("ds_actor", self.sign({"a": actor, "e": expiry}, "actor"))`
- **Signing**: itsdangerous `URLSafeSerializer(self._secret, namespace)` β€” `_secret` is a random 32-byte hex per process unless supplied (app.py:431). Namespaces like `"actor"`/`"messages"` keep cookie types separate.
- **Verifying**: the built-in hook impl unsigns the cookie, checks for an `"a"` key, and decodes the base62-encoded `"e"` expiry timestamp. `BadSignature` β†’ anonymous (`None`).

So plugins like *datasette-auth-cookies* / *-auth-github* implement a login page that sets this cookie; the core handles the rest, including `/-/logout` (which just clears it via `delete_actor_cookie`).

## 3. API tokens (`datasette/tokens.py`)

A **token handler** system: plugins register handlers via `register_token_handler`; the core provides `Datasette.create_token()` / `verify_token()` which try each registered handler in order and return an actor dict for the first that recognizes the token. A client presents a token (e.g. *datasette-auth-tokens* reads the `X-Auth-Token` header) inside its own `actor_from_request` impl, calling `verify_token()`. Tokens can carry `TokenRestrictions` limiting which actions/databases/tables they may use.

## 4. The root actor (`--root`)

The only "built-in" login: with `--root`, an actor of `{"id": "root"}` gets a global allow-all rule for every action (including writes and debug), still overridable by explicit deny rules in `datasette.yaml`. It's meant for local dev.

## 5. Permissions (separate from auth)

Once you have an actor, checks go through `datasette.allowed(action=..., resource=..., actor=...)` β†’ `allowed_many()` (app.py ~1881). Rules come from plugins' `check_permission` hooks and YAML config, are stored in the internal SQLite DB, resolved with a single SQL query per batch, cached per request via a contextvar (`permissions.py`). Actor display names resolve from `display`/`name`/`username`/`login`/`id`.

**In short**: plugins decide *who* (via `actor_from_request`, using signed cookies or tokens); Datasette core decides *what they can do* (permission rules + `allowed()`), with the actor dict flowing through everything as `request.actor`. Full docs in `docs/authentication.rst`.
=== isError present? ===
11
🧠 Thinking

The output looks good. Since isError appears 11 times within the JSONL, let's check how it's structured and verify that my error handling works correctly. I'll take a look at one of those lines.

πŸ”§ bash

{
  "command": "cd /private/tmp/transform && python3 -c \"\nimport json\nfor line in open('2026-08-15T19-10-16-083Z_01a006d4-fa13-76b4-83f3-1a0ab01392c5.jsonl'):\n    d = json.loads(line)\n    if 'isError' in line and d.get('type')=='message':\n        m=d['message']\n        print(m.get('role'), '| isError=', m.get('isError'), '| content types:', [c.get('type') for c in m.get('content',[])][:3])\n\" | head -5 && grep -n '❌' transcript.md | head -3"
}

↩️ bash result

toolResult | isError= False | content types: ['text']
toolResult | isError= False | content types: ['text']
toolResult | isError= False | content types: ['text']
toolResult | isError= False | content types: ['text']
toolResult | isError= False | content types: ['text']
🧠 Thinking

All the tool results in this file have isError: false at the message level, and my code checks block.get("isError"), but here block is the message dict itself (the fallback path for src), so it works β€” in fact, when isError is False it displays as "result", which is correct. If isError were True, it should display as ❌ failed.

Everything has been verified to work. The script handles the following:

  • session/model_change/thinking_level_change β†’ header metadata
  • user messages β†’ ## πŸ’¬ User section
  • assistant thinking β†’ collapsible <details> (can be omitted with --no-thinking)
  • toolCall β†’ fenced code block of JSON arguments
  • toolResult β†’ labeled fenced output, error detection via isError

Done. Wrapping up concisely.

Done. The converter is at jsonl_to_md.py and has been tested on your session file (output: transcript.md, 39.5KB).

Usage:

python3 jsonl_to_md.py input.jsonl [-o output.md] [--no-thinking]

What it handles (verified against your JSONL's actual structure):

Input Markdown output
session, model_change, thinking_level_change lines Header with session ID, model (spark/qwen3.8-27b), cwd, thinking level
user message ## πŸ’¬ User section
Assistant thinking blocks Collapsible <details>🧠 Thinking</details> (strip with --no-thinking)
toolCall blocks Labeled fenced block: ```bash for bash, JSON args for others
toolResult **↩️ tool result** / **❌ failed** (checks isError) + fenced output

Edge cases covered: fences auto-widen to 4+ backticks if content contains triple-backtick code blocks, invalid JSONL lines are skipped with a warning instead of crashing, and multi-tool-call assistant turns render in order.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment