Sitelet https://github.com/microsoft/TypeAgent/pull/3096
Skip to content

Support local embeddings in KnowPro; configurable embedding size and batch size - #3096

Merged
Dominic Nguyen (datduyng) merged 2 commits into
mainfrom
domnguyen/knowpro-local-embedding
Sep 29, 2026
Merged

Dominic Nguyen (datduyng) merged 2 commits into
mainfrom
domnguyen/knowpro-local-embedding

Conversation

@datduyng

Copy link
Copy Markdown
Contributor

Adds support for local embeddings in KnowPro, and makes the embedding size and batch size configurable:

embedding:
  provider: local
  model: Xenova/all-MiniLM-L6-v2
  size: 384
  maxBatchSize: 32
  • New embedding.size / embedding.maxBatchSize config (TYPEAGENT_EMBEDDING_SIZE / _MAX_BATCH_SIZE).
  • KnowPro sizes its indexes with getEmbeddingSize() (384 for local, 1536 for hosted) instead of a hardcoded 1536. Caller-supplied models keep the old default.
  • OpenAI/Azure/Copilot now use the configured model, size, and maxBatchSize too.

Add embedding.size and embedding.maxBatchSize to the typed config (TYPEAGENT_EMBEDDING_SIZE / _MAX_BATCH_SIZE). KnowPro now sizes indexes from getEmbeddingSize() instead of a hardcoded 1536, so the local MiniLM provider (384-d) works out of the box.
@datduyng
Dominic Nguyen (datduyng) added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit 3420a06 Sep 29, 2026
27 checks passed
const MaxBatchSize = 64;
const DefaultMaxBatchSize = 64;

export type CopilotEmbeddingOptions = {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why not just call this EmbeddingOptions?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

disregard...thought I was looking at embeddingProvider.ts.

const LocalDefaultEmbeddingSize = 384; // Xenova/all-MiniLM-L6-v2
const HostedDefaultEmbeddingSize = 1536; // ada-002 / text-embedding-3-small

function readPositiveInt(name: string): number | undefined {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be useful to warn if this reads a non-positive integer here. Right now this just swallows the misconfiguration and just continues so no one's the wiser.

/**
* The embedding vector size for the configured provider. An explicit
* `embedding.size` (`TYPEAGENT_EMBEDDING_SIZE`) always wins; otherwise the
* default model's size is used (set `size` when using a non-default model) (384 for the local MiniLM model, 1536 for

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

THis comment and the values above can drift. I recommend reusing the constant names here so if their values change the comment is still valid.

});
case "copilot":
return createCopilotEmbeddingModel(
process.env[EmbeddingEnvVars.MODEL]?.trim() ||

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We moved away from .env config a while ago but there are still places where we go directly to env vars. We should avoid adding new ones and instead should work to get rid of existing process.env references. Everything should migrate to the YAML based configuration: https://github.com/microsoft/TypeAgent/tree/main/ts/packages/config.

These new configuration options should go into the config package and that can do the proper process.env, .env, or YAML loading and then this package can just use the values.

? {}
: {
model: settings.modelName,
model: options?.modelName ?? settings.modelName,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Here you're overriding a pool setting with the supplied options...maybe another good place to make a n informational logging call so there's enough trace information to troubleshoot issues if needed.

embeddingModel = tryCreateEmbeddingModel();
embeddingSize ??= getEmbeddingSize();
}
embeddingSize ??= 1536;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should modify this line to use the configured default from aiclient.embeddingProvider.

jebrans pushed a commit to jebrans/TypeAgent that referenced this pull request Sep 30, 2026
Follow-up to microsoft#3096.

- `embeddingProvider` reads the `embedding:` section from the typed
runtime config instead of `process.env`.
- Config warns when a positive-integer setting (e.g. `embedding.size`)
is invalid, instead of ignoring it silently.
- `getEmbeddingSize` doc refers to the default-size constants, so it
cannot drift.
- `createEmbeddingModel` logs model/batch-size overrides of the pool
settings (`typeagent:openai`).
- KnowPro uses `getEmbeddingSize()` for caller-supplied models too,
instead of a hardcoded 1536.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants