Skip to content

fix: use a realistic fallback divisor for collector token estimates - #6672

Merged
timothycarambat merged 2 commits into
Mintplex-Labs:masterfrom
moneebullah25:fix/collector-token-estimate-monotonic
Oct 9, 2026
Merged

timothycarambat merged 2 commits into
Mintplex-Labs:masterfrom
moneebullah25:fix/collector-token-estimate-monotonic

Conversation

@moneebullah25

@moneebullah25 moneebullah25 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request Type

  • ✨ feat (New feature)
  • 🐛 fix (Bug fix)
  • ♻️ refactor (Code refactoring without changing behavior)
  • 💄 style (UI style changes)
  • 🔨 chore (Build, CI, maintenance)
  • 📝 docs (Documentation updates)

Relevant Issues

resolves #6668

Description

TikTokenTokenizer.DIVISOR in collector/utils/tokenizer/index.js goes from 8 to 4. The "too long" guard is untouched (it still protects the CPU); only the fallback estimate it returns changes.

With the old divisor of 8, token_count_estimate dropped sharply at exactly 5120 chars (e.g. 1003 -> 640 tokens for English prose), so the estimate was not monotonic and under-counted roughly 2x for typical English. With 4 the same input gives 1280, close to the real count and no longer a cliff.

Tests: added collector/__tests__/utils/tokenizer/index.test.js:

  • 5119 -> 5120 chars does not decrease
  • estimate at 5120 chars stays within 0.8x-1.5x of the real encoding
  • exact ceil(length / 4) once the guard trips
  • empty input returns 0

Two of these fail on the old divisor of 8. The full collector suite passes (14 suites, 410 tests); eslint and prettier --check are clean on the changed files.

Additional Information

CJK / JSON / whitespace-heavy text above the guard is still under-counted by a fixed divisor; that is out of scope here (related only: #6142).

Developer Validations

  • I ran yarn lint from the root of the repo & committed changes
  • Relevant documentation has been updated (if applicable)
  • I have tested my code functionality
  • Docker build succeeds locally

@shatfield4 shatfield4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Text above 5120 chars gets an estimated token count of length / DIVISOR. Real text averages about 4 chars per token, so /8 under-counted long docs by about half, and they passed the context window check when they shouldn't have. /4 is close to real counts and only overestimates, which is the safe direction. I removed the test since it only held for its sample text.

@shatfield4 shatfield4 added the PR:Ready-to-merge PR has been reviewed by core team and is ready to merge label Oct 8, 2026
@timothycarambat
timothycarambat merged commit 3034417 into Mintplex-Labs:master Oct 9, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

PR:Ready-to-merge PR has been reviewed by core team and is ready to merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: token_count_estimate drops ~35-45% at 5120 chars (non-monotonic; /8 fallback under-counts)

3 participants