-
Notifications
You must be signed in to change notification settings - Fork 21
Permalink
Choose a base ref
{{ refName }}
default
Choose a head ref
{{ refName }}
default
Comparing changes
Choose two branches to see what’s changed or to start a new pull request.
If you need to, you can also or
learn more about diff comparisons.
Open a pull request
Create a new pull request by comparing changes across two branches. If you need to, you can also .
Learn more about diff comparisons here.
base repository: sqliteai/sqlite-ai
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: 1.0.5
Could not load branches
Nothing to show
Loading
Could not load tags
Nothing to show
{{ refName }}
default
Loading
...
head repository: sqliteai/sqlite-ai
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: 1.0.6
Could not load branches
Nothing to show
Loading
Could not load tags
Nothing to show
{{ refName }}
default
Loading
- 3 commits
- 5 files changed
- 3 contributors
Commits on Aug 24, 2026
-
fix(chat): keep the chat usable when it runs out of context (#30)
Running a chat out of context did not just fail the turn, it corrupted the conversation: every later turn failed too, and the failure got worse each time. turn 11: Context size exceeded (256, 266) turn 12: Context size exceeded (256, 276) turn 13: Context size exceeded (256, 286) turn 14: Context size exceeded (256, 296) llm_chat_run() appends the user message to the history before generating, then bails straight out of the token loop on failure - never reaching llm_chat_save_response(). The user turn is left stranded with no assistant reply and prev_len is never advanced, so the template delta for the next turn re-includes it. Hence the number climbing by a turn's worth of tokens on every retry, and a chat that can never be used again. Commit the partial turn before reporting the failure. The streaming cursor path already does this in xClose for the same reason; the non-streaming path disagreeing with it was the bug. The same run now reports a stable Context size exceeded (256, 267) on every attempt, and the saved history has an assistant row for every user row instead of trailing orphans (28 rows vs 24). Also fixes the guard that decides this. llama_memory_seq_pos_max() returns the highest position in the sequence, so occupancy is that + 1; treating it as a count let exactly one over-large batch through for llama_decode() to reject with "could not find a KV slot" instead - which is why the failure used to surface from the decode rather than from the guard meant to prevent it. That is the 266 -> 267 difference. The comparison also promoted int32_t to uint32_t, so an empty cache (seq_pos_max == -1) with a zero-token batch tripped it falsely. This is the recoverable half of #28. The turn still returns an error rather than the partial reply plus a stop reason; that part is deliberately left open, since it changes the contract and deserves its own decision. test_chat_context_full_is_recoverable talks until the context fills, then keeps going, and asserts the guard reports it, that the requirement does not grow across retries, and that the history does not end on a user turn with no reply. It compares only from the second failure onward: the first can land mid-generation, where the batch is a single token, so it legitimately reports a different number from the prompt-batch retries that follow. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012KRojTrn4Q4SaZQpqcwpAtConfiguration menu - View commit details
-
Copy full SHA for d2b4a63 - Browse repository at this point
Copy the full SHA d2b4a63View commit details -
fix(context): honour an explicit context_size of 512 (#29)
llm_context_create*() decided whether the caller had asked for a context size by comparing the parsed value against llama_context_default_params().n_ctx: struct llama_context_params defaults = llama_context_default_params(); if (ai->model && ctx_params.n_ctx == defaults.n_ctx) { ctx_params.n_ctx = 0; // 0 = use the model's training window } That default is 512 (llama-context.cpp:2772), a value a caller can perfectly well pass, so an explicit context_size=512 was indistinguishable from unset and was silently replaced by the model's full training window: context_size=64 -> n_ctx=256 context_size=512 -> n_ctx=32768 <-- asked for 512, got 64x that context_size=513 -> n_ctx=768 n_ctx=512 hit it too, since the check looked at the resolved value rather than at which key was written. The intent was right, only the detection was wrong. llama.cpp already defines n_ctx = 0 as "use the training window", so start from that sentinel instead of inferring it after the fact: the caller's value now always survives, and 0 keeps its documented meaning. Also stops context_size=0 driving n_batch to 0, which llama will not accept. Behaviour change, and a quiet one: anyone passing context_size=512 or n_ctx=512 was getting the model's whole window and now gets 512. Nothing errors - the context simply becomes what was asked for, so a conversation that used to fit may now reach the limit. That is the point: silently ignoring the configuration is the bug. #30 makes reaching the limit recoverable rather than fatal to the chat; the remaining half of #28 - returning the partial reply plus a stop reason instead of an error - is still open. API.md said llm_context_create_chat() and llm_context_create_textgen() were equivalent to context_size=4096. Their presets are empty, so both inherit the model's training window; documented as such, along with 0 on context_size/n_ctx. test_context_size_is_honoured covers 256/512/1024 exactly (llama pads n_ctx to a multiple of 256), both spellings, and that omitting the key - or passing 0 - still auto-sizes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012KRojTrn4Q4SaZQpqcwpAtConfiguration menu - View commit details
-
Copy full SHA for a5d04a2 - Browse repository at this point
Copy the full SHA a5d04a2View commit details -
Configuration menu - View commit details
-
Copy full SHA for ec8992e - Browse repository at this point
Copy the full SHA ec8992eView commit details
Loading
This comparison is taking too long to generate.
Unfortunately it looks like we can’t render this comparison for you right now. It might be too big, or there might be something weird with your repository.
You can try running this command locally to see the comparison on your machine:
git diff 1.0.5...1.0.6