Sitelet https://github.com/sqliteai/sqlite-ai/compare/1.0.5...1.0.6
Skip to content
Permalink

Comparing changes

Choose two branches to see what’s changed or to start a new pull request. If you need to, you can also or learn more about diff comparisons.

Open a pull request

Create a new pull request by comparing changes across two branches. If you need to, you can also . Learn more about diff comparisons here.
base repository: sqliteai/sqlite-ai
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: 1.0.5
Choose a base ref
...
head repository: sqliteai/sqlite-ai
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: 1.0.6
Choose a head ref
  • 3 commits
  • 5 files changed
  • 3 contributors

Commits on Aug 24, 2026

  1. fix(chat): keep the chat usable when it runs out of context (#30)

    Running a chat out of context did not just fail the turn, it corrupted the
    conversation: every later turn failed too, and the failure got worse each time.
    
        turn 11: Context size exceeded (256, 266)
        turn 12: Context size exceeded (256, 276)
        turn 13: Context size exceeded (256, 286)
        turn 14: Context size exceeded (256, 296)
    
    llm_chat_run() appends the user message to the history before generating, then
    bails straight out of the token loop on failure - never reaching
    llm_chat_save_response(). The user turn is left stranded with no assistant reply
    and prev_len is never advanced, so the template delta for the next turn
    re-includes it. Hence the number climbing by a turn's worth of tokens on every
    retry, and a chat that can never be used again.
    
    Commit the partial turn before reporting the failure. The streaming cursor path
    already does this in xClose for the same reason; the non-streaming path
    disagreeing with it was the bug. The same run now reports a stable
    
        Context size exceeded (256, 267)
    
    on every attempt, and the saved history has an assistant row for every user row
    instead of trailing orphans (28 rows vs 24).
    
    Also fixes the guard that decides this. llama_memory_seq_pos_max() returns the
    highest position in the sequence, so occupancy is that + 1; treating it as a
    count let exactly one over-large batch through for llama_decode() to reject with
    "could not find a KV slot" instead - which is why the failure used to surface
    from the decode rather than from the guard meant to prevent it. That is the
    266 -> 267 difference. The comparison also promoted int32_t to uint32_t, so an
    empty cache (seq_pos_max == -1) with a zero-token batch tripped it falsely.
    
    This is the recoverable half of #28. The turn still returns an error rather than
    the partial reply plus a stop reason; that part is deliberately left open, since
    it changes the contract and deserves its own decision.
    
    test_chat_context_full_is_recoverable talks until the context fills, then keeps
    going, and asserts the guard reports it, that the requirement does not grow
    across retries, and that the history does not end on a user turn with no reply.
    It compares only from the second failure onward: the first can land
    mid-generation, where the batch is a single token, so it legitimately reports a
    different number from the prompt-batch retries that follow.
    
    Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_012KRojTrn4Q4SaZQpqcwpAt
    andinux and claude authored Aug 24, 2026
    Configuration menu
    Copy the full SHA
    d2b4a63 View commit details
    Browse the repository at this point in the history
  2. fix(context): honour an explicit context_size of 512 (#29)

    llm_context_create*() decided whether the caller had asked for a context size by
    comparing the parsed value against llama_context_default_params().n_ctx:
    
        struct llama_context_params defaults = llama_context_default_params();
        if (ai->model && ctx_params.n_ctx == defaults.n_ctx) {
            ctx_params.n_ctx = 0;   // 0 = use the model's training window
        }
    
    That default is 512 (llama-context.cpp:2772), a value a caller can perfectly
    well pass, so an explicit context_size=512 was indistinguishable from unset and
    was silently replaced by the model's full training window:
    
        context_size=64   -> n_ctx=256
        context_size=512  -> n_ctx=32768     <-- asked for 512, got 64x that
        context_size=513  -> n_ctx=768
    
    n_ctx=512 hit it too, since the check looked at the resolved value rather than
    at which key was written.
    
    The intent was right, only the detection was wrong. llama.cpp already defines
    n_ctx = 0 as "use the training window", so start from that sentinel instead of
    inferring it after the fact: the caller's value now always survives, and 0 keeps
    its documented meaning. Also stops context_size=0 driving n_batch to 0, which
    llama will not accept.
    
    Behaviour change, and a quiet one: anyone passing context_size=512 or n_ctx=512
    was getting the model's whole window and now gets 512. Nothing errors - the
    context simply becomes what was asked for, so a conversation that used to fit
    may now reach the limit. That is the point: silently ignoring the configuration
    is the bug. #30 makes reaching the limit recoverable rather than fatal to the
    chat; the remaining half of #28 - returning the partial reply plus a stop reason
    instead of an error - is still open.
    
    API.md said llm_context_create_chat() and llm_context_create_textgen() were
    equivalent to context_size=4096. Their presets are empty, so both inherit the
    model's training window; documented as such, along with 0 on context_size/n_ctx.
    
    test_context_size_is_honoured covers 256/512/1024 exactly (llama pads n_ctx to a
    multiple of 256), both spellings, and that omitting the key - or passing 0 -
    still auto-sizes.
    
    Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_012KRojTrn4Q4SaZQpqcwpAt
    andinux and claude authored Aug 24, 2026
    Configuration menu
    Copy the full SHA
    a5d04a2 View commit details
    Browse the repository at this point in the history
  3. Configuration menu
    Copy the full SHA
    ec8992e View commit details
    Browse the repository at this point in the history
Loading