“There is a number that has quietly become one of the most important line items in every serious AI engineering budget: the token count. Every call to AI LLM is billed by the token. A single token is roughly four characters of English text. That sounds benign until you run your first agentic workflow at scale and discover that a single SRE incident investigation can consume 65,000 input tokens in one session just in log data, stack traces, and tool outputs.” — Headroom introduction
A new approach in context engineering is now emerging: compress what you send before it reaches the model. Process raw, verbose tool outputs before calling to AI LLM, strip out redundancy, and ship only the semantic effective signal to the LLM.
Context semantic compression (Prompt Context) condenses lengthy conversation histories, documents and prompts into compact tokens or summarized text for AI model input. Its core strength lies in resource efficiency: it drastically cuts token consumption, lowering API costs and bypassing rigid context window limits to process far more background data in one request. Faster inference speed also emerges, as smaller input batches reduce model computation latency for real-time chat and retrieval-augmented generation.
Yet aggressive compression also introduces irreversible information loss; once fine-grained context is discarded, the model cannot revisit critical background details to refine answers, hurting complex reasoning tasks like legal analysis or mathematical problem-solving.
The suitable choice of compression rate is very important. Compression rate is a metric to measure how much original data is shrunk after compression. It quantifies the space-saving efficiency of compression algorithms, compression rate = compressed size / original size. If compression rate ~ 0.5, means the compressed size is roughly 1/2 of the original text size.
This is the philosophy behind Tokenflow — a context semantic compression API built for AI engineers who are serious about cost control without sacrificing answer quality.
Tokenflow is a context semantic compression API for AI agents and LLM-powered applications. Its core promise: 50–90% fewer tokens, almost same answers.
It compresses everything an AI agent reads — tool outputs, log files, RAG retrieval chunks, source code, conversation history, and structured JSON — before that content reaches the LLM provider. The LLM receives a semantically equivalent but dramatically smaller representation, processes it correctly, and answers as accurately as it would have with the full payload.
Tokenflow is not a model. It is not a provider. It is a middleware layer — an API processing pipeline that sits between your application and the cloud LLM.
A traditional chatbot sends a question and gets an answer. The context is small. AI agents operate on an entirely different paradigm. A coding agent in a single session might:
Each step appends a tool output to the conversation context, which is sent in full with the next step. Finally, the agent may be carrying 150,000 tokens of accumulated context — most of it raw, unprocessed noise.
Our context semantic compression engine is a kind of lossy semantic compression algorithm. It is a very fast sampling algorithm in dynamically distributed semantic space.
Calling Tokenflow API is very easy. Here is a Node.js example (download here or https://github.com/). First of all:
npm install @babel/core @babel/parser @babel/traverse @babel/types @types/node htmlparser2 marked openai
Most Node libraries are needed in index.js for text and code segmentation, such as processSegment, which is an easy, default segment tool. But you can develop more sophisticated segmentation tools for text, js, ts, c, c++, cpp, java, go, or python.
The parameters of Tokenflow API are model, compression rate, and input. The model parameter is language selection. Right now it supports only five major international languages: English, French, German, Spanish, Italian. For example, model: 'English'. Other languages are on the way. For programming code, use model: 'code'. The compression rate (CR) is in range 0 < CR < 1.0. CR = 0.5–0.1 means 50–90% fewer tokens. The input format is a data group, such as [ { content: string }, { content: string }, … ].
import OpenAI from "openai";
import fs from "fs";
import path from "path";
import { processSegment } from "./index.js";
const client = new OpenAI({
baseURL: 'http://www.webtokenflow.com:3000/v1',
apiKey: 'my-secret-1234567891'
});
async function txtParser() {
// example for txt
const txtStr = fs.createReadStream("./input.txt");
const txtResult = await processSegment(txtStr, "txt");
// example for md
// const mdSource = fs.createReadStream("./test.md");
// const mdResult = await processSegment(mdSource, "md");
// call api
const res = await client.chat.completions.create({
model: 'English',
compressionrate: 0.1,
input: JSON.stringify(txtResult)
});
// return result
const results = JSON.parse(res.choices[0].message.content);
// write result into a file
for (const result of results) {
fs.appendFileSync('./sampled.txt', result.content + '\n', 'utf8');
}
console.log(res.usage);
}
txtParser().catch(console.error);
Index.js is only for text and code segmentation.
Text segmentation is an NLP technique that divides unstructured text into semantic sub-units via rule, statistics or AI models. In an AI context window, if the size of request is large, for example, 1000+ paragraphs, it should use paragraph as natural element by line break (\n). However, if the size of request is roughly < 100 paragraphs, the sentence as segment element is a better choice. In very few paragraphs, words as basic semantic elements are much better; in this case, it is equivalent to keyword abstraction in NLP.
Code segmentation is a specialized, syntax-dependent branch of text segmentation, with four mainstream implementation schemes: simple delimiter splitting, bracket scope matching, AST/Tree-sitter syntax parsing, and comment marker segmentation.
Due to divergent block syntax across programming languages, industrial-grade code segmentation solutions usually adopt Tree-sitter AST parsing as the core foundation, supplemented by bracket balance logic, to achieve robust, multi-language compatible chunk splitting for large-scale text and source code processing pipelines.