Processing 10,000 lines of code (~112,000 tokens) with Text Embedding 004 costs $0.00224 for codebase ingestion and $0.00224 for an AI-powered code review and refactoring pass.
Feed 10,000 lines of code into the prompt context for repository search, Q&A, or architecture planning.
Ingest 10,000 lines of code and generate audit findings, unit test recommendations, and refactor diffs.
Text Embedding 004 writes 10,000 lines of code from scratch based on product specifications.
| Model | Provider | Code Ingestion | Cached Ingestion | Code Review Cost | Context Limit |
|---|---|---|---|---|---|
| Text Embedding 004 (Current) | $0.00224 | $0.00056 | $0.00224 | 8,192 | |
| Text Embedding 3 (Small) | openai | $0.00224 | $0.00056 | $0.00224 | 8,191 |
| Text Embedding 3 (Large) | openai | $0.0146 | $0.00364 | $0.0146 | 8,191 |
On average, code yields approximately 11.2 tokens per line in Text Embedding 004 (sentencepiece tokenizer). Indentation, brackets, camelCase variable names, and comments slightly increase token density compared to plain English text. 10,000 lines of code produces approximately 112,000 tokens.
Sending 10,000 lines of code as context and generating a thorough code review with recommendations costs approximately $0.00224. Utilizing prompt caching on repeat turns or static repository definitions drops this to $0.00056.
Text Embedding 004 has a context window of 8,192 tokens. 10,000 lines of code consumes 1367.19% of its total available context.