llm-chat-completions-server 0.1a0

Simon Willison's Weblog Simon Willison's Weblog

Release: https://github.com/simonw/llm-chat-completions-server/releases/tag/0.1a0">llm-chat-completions-server 0.1a0


A key goal of the new content-addressable logs https://simonwillison.net/2026/Jul/30/llm-rc1/">in LLM 0.32rc1 was being able to support OpenAI Chat Completion style requests where each incoming message extends the previous conversation, like this:


curl http://localhost:8002/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3.5-4b",
"messages": [
{"role": "user", "content": "Capital of France?"},
{"role": "assistant", "content": "Paris."},
{"role": "user", "content": "Germany?"}
]
}'

Here the conversation state is tracked by the client, so each of these requests gets longer and longer.

The new schema design in LLM is designed to de-duplicate these using hashes of the individual message parts.


To test that out, I built this plugin:


uv tool install llm --pre
llm install llm-chat-completions-server
llm chat-completions-server -p 9001

Running this starts a localhost server on port 9001 that exposes your full collection of LLM models (from any plugins you have installed) using a ChatGPT Completions compatible endpoint.


GPT-5.6 Sol https://gist.github.com/simonw/53be513c1bd4a29a7aa480d9bde9b4a5">wrote the whole thing - it turns out it knows the OpenAI Chat Completions API shape really well.




Tags: https://simonwillison.net/tags/projects">projects, https://simonwillison.net/tags/openai">openai, https://simonwillison.net/tags/llm">llm

Read full article at Simon Willison's Weblog →