Blockify Benchmark: The Proprietary RAG Accuracy Data, With Full Methodology
Every Blockify performance figure in one place — vector-search accuracy, token efficiency, data reduction, aggregate accuracy, and healthcare accuracy — each traced to a documented benchmark with named sources. Cite it, replicate it, scrutinize it.
What does the Blockify benchmark show? In documented enterprise benchmarks, Blockify improves RAG accuracy by up to 78X on an aggregate basis, delivers 2.29X more accurate vector search (a 56.26% precision gain), reduces token usage per query by 3.09X, and shrinks datasets by up to 40X to roughly 2.5% of their original size. In a safety-critical healthcare evaluation across nine clinical questions, Blockify improved combined accuracy and source fidelity by 261.11% on average, peaking at 650% on diabetic-ketoacidosis management. Every figure below is traced to its source benchmark and method.
How These Benchmarks Were Measured
These are proprietary enterprise benchmarks, and we publish the method alongside every number so you can scrutinize or replicate it. The headline figures come from two documented studies: the Blockify Performance Analysis conducted on a Big Four consulting firm’s dataset (17 documents / 298 pages), and an “Evaluation of Blockify” medical-accuracy study across nine clinical questions. Both compare Blockify’s distilled IdeaBlocks against naive fixed-size (~1,000-character) chunking — a method, not a named product. The comparison is scoped to retrieval quality — the same property that LLM evaluation metrics such as faithfulness and context precision/recall are designed to capture.
The most important distinction on this page is between measured single-benchmark figures and documented cross-deployment averages. Vector-search accuracy (2.29X) and token efficiency (3.09X) are measured directly in the Big Four benchmark. The aggregate accuracy and data-reduction headlines you may have seen elsewhere — “up to 78X” and “up to 40X” — are cross-deployment averages. On this single Big Four dataset, the transparent worked example is a 68.44X aggregate and a 2.00X base (29.93X enterprise-adjusted) word reduction. We label every number accordingly. Each gain measured here comes from how the source data is represented rather than from switching to a different model — the same lens Iternal applies when reviewing the most accurate enterprise AI platforms.
Vector accuracy is measured as the mean cosine distance to the best-match result — the smallest distance — which deliberately favors the chunking baseline. Token counts assume a 4:1 character-to-token ratio. Enterprise-scale figures apply IDC’s documented 15:1 average data-duplication factor (a range of 8:1 to 22:1). Where a figure is an average across deployments rather than this benchmark’s direct output, it is marked as such.
Vector-Search Accuracy: 2.29X
Blockify IdeaBlocks return 2.29X more accurate vector-search results than naive 1,000-character chunking (a 56.26% precision improvement).
Source: Iternal Technologies Blockify Performance Analysis (Big Four consulting-firm dataset, 17 documents / 298 pages).
Method: mean cosine distance to best-match, 0.3624 chunking vs 0.1585 distilled IdeaBlocks; best match defined as smallest distance to deliberately favor the chunking baseline.
Vector search retrieves the passage whose embedding sits closest to the query. Smaller cosine distance means a tighter, more relevant match. Distilling documents into IdeaBlocks cuts the average best-match distance by more than half, so the model grounds its answer on the right passage far more often.
| Representation | Avg. cosine distance to best match | Relative accuracy |
|---|---|---|
| Legacy chunking (~1,000 chars) | 0.3624 | Baseline |
| IdeaBlocks (undistilled) | 0.1833 | 1.98X |
| IdeaBlocks (distilled) | 0.1585 | 2.29X |
See the full vector-accuracy methodology (Performance Analysis §09)
Token Efficiency: 3.09X
Blockify cuts tokens processed per query by 3.09X, an estimated $738,000/year saving at scale.
Source: Iternal Technologies Blockify Performance Analysis.
Method: ~303 tokens/chunk vs ~98 tokens/distilled block across top-5 retrieval; priced at $0.72 per 1M tokens (Llama 3.3 70B) over 1 billion annual queries.
Because IdeaBlocks are compact and deduplicated, each retrieval pulls far fewer tokens into the context window. Fewer tokens per query compounds into lower inference cost, lower latency, and more headroom under a fixed context budget.
| Measure | Legacy chunking | Distilled IdeaBlocks | Improvement |
|---|---|---|---|
| Avg. tokens per unit retrieved | ~303 | ~98 | 3.09X |
| Estimated tokens / year | ~1.515T | ~490B | 3.09X |
| Estimated annual saving | — | — | $738,000/yr |
See the token-efficiency worked example (Performance Analysis §08)
Data Reduction: Up to 40X (~2.5% of Original Size)
Blockify distills enterprise corpora to roughly 2.5% of their original size (up to 40X smaller) on average.
Source: Iternal Technologies Blockify Performance Analysis.
Method: cross-deployment average; the Big Four dataset measured a 2.00X base word reduction (88,877 → 44,537 words), rising to 29.93X once the IDC 15:1 enterprise data-duplication factor is applied.
“Up to 40X” is a documented cross-deployment average, not this benchmark’s measured output. The transparent worked example on the Big Four dataset is a 2.00X base word reduction (88,877 → 44,537 words), which rises to 29.93X by word count (and 27.82X by character count) once IDC’s 15:1 enterprise duplication factor is applied. Smaller datasets mean fewer tokens, lower cost, and a golden corpus small enough for humans to review.
See the data-reduction table (Performance Analysis §08, Table 7)
Aggregate Accuracy: Up to 78X
Blockify delivers up to 78X aggregate RAG accuracy improvement across enterprise deployments.
Source: Iternal Technologies Blockify Performance Analysis.
Method: aggregate of vector accuracy and data-volume reduction; the Big Four benchmark measured 68.44X (4.56X base × ~15X IDC enterprise duplication factor). 78X is the documented cross-deployment average, not a single-test result.
78X is the documented cross-deployment average. On this single Big Four benchmark, the transparent measured aggregate is 68.44X. The derivation is fully shown: a 4.56X base improvement (2.29X vector accuracy × 2.00X word-count reduction) is multiplied by roughly 15X — IDC’s average enterprise data-duplication factor (8:1 to 22:1) — for a 68.44X enterprise aggregate. We present 68.44X as the worked example precisely because it is what this dataset produced; 78X is the average across deployments.
Healthcare Accuracy: Up to 650% (261.11% Average)
In a healthcare RAG evaluation, Blockify improved accuracy and source fidelity by 261.11% on average and up to 650% on the highest-stakes topic (diabetic-ketoacidosis management).
Source: Iternal Technologies “Evaluation of Blockify” medical accuracy study.
Method: 9 clinical questions, Blockify vs chunking; 650% is the peak single-topic gain (DKA), 261.11% is the mean across all nine.
650% is the peak, not the average. Across nine clinical questions, Blockify improved combined accuracy and source fidelity by 261.11% on average, with the largest gains on the highest-stakes topics. In one safety-critical case, chunking recommended “D5W” (a dextrose solution) as an initial IV fluid, where standard protocol calls for isotonic saline first; the Blockify response avoided the error by not specifying the fluid type prematurely.
| Clinical question | Accuracy & source-fidelity improvement |
|---|---|
| DKA management (peak) | 650% |
| Pneumonia lab tests | 500% |
| Headache red flags | 250% |
| Heart-failure prognosis | 250% |
| Other queries | 100–300% |
| Average across 9 questions | 261.11% |
How to Evaluate RAG: Metrics, Frameworks, and What the Numbers Mean
RAG evaluation measures a retrieval-augmented generation system in two halves: retrieval quality, scored with recall@k, MRR and nDCG on benchmarks such as BEIR, and generation quality, scored with faithfulness, answer relevance and context precision. Libraries including RAGAS, TruLens, DeepEval, ARES and Arize Phoenix automate both halves.
Every figure on this page is a RAG evaluation result. The vocabulary below is the one practitioners use to describe results like them, so the 2.29X vector-search gain and the 261.11% healthcare gain can be read against a recognized scorecard rather than taken on trust — and so the same comparison can be reproduced on a different corpus.
Retrieval Metrics: Did the Right Passage Come Back?
Retrieval is scored on the ranked list of passages a query returns, before the model writes anything. Five measures do most of the work, and which one leads depends on the failure the system cannot afford.
| Metric | What it answers | Lead with it when |
|---|---|---|
| Recall@k | Of the passages that should have been retrieved, how many appear in the top k? | A handful of passages in the corpus can answer the question and missing one is the failure that matters. |
| Precision@k | Of the k passages returned, how many are actually relevant? | The context window is tight and irrelevant passages crowd out the correct one. |
| MRR (mean reciprocal rank) | How high in the list does the first correct passage sit? | The model reads the top result far more closely than the rest of the list. |
| nDCG@k | Are the most relevant passages ranked highest, discounted by position? | Relevance is graded rather than yes-or-no, and rank order changes the answer. |
| Hit rate@k | Did any correct passage come back at all? | A coarse smoke test before the finer metrics are worth computing. |
For a portable read on a retriever or an embedding model, teams run BEIR, the heterogeneous retrieval benchmark published by Thakur et al. (NeurIPS 2021 Datasets and Benchmarks), which bundles 18 public datasets across nine retrieval tasks and scores zero-shot generalization, conventionally at nDCG@10. BEIR answers whether a retriever is strong in general. It does not answer whether it is strong on your documents, which is why an enterprise program labels a query set from its own corpus as well.
Generation Metrics: Was the Answer Grounded in What Came Back?
Retrieval can succeed and the answer still be wrong. The generation half of the evaluation scores the response against the passages that were actually retrieved. The working definitions most teams use were published with the RAGAS project (Es et al., 2023).
| Metric | What it answers | What it catches |
|---|---|---|
| Faithfulness | Is every claim in the answer supported by the retrieved passages? | Catches hallucination: fluent text with nothing behind it. |
| Answer relevance | Does the response address the question that was actually asked? | Catches on-topic padding and evasive answers. |
| Context precision | Are the passages the answer relied on ranked at the top of the context? | Catches a right answer reached in spite of noisy retrieval. |
| Context recall | Does the retrieved context contain everything a reference answer needs? | Catches partial answers caused by one missing passage. |
| Answer correctness and source fidelity | Does the answer match the reference, and does it point at the right source? | Catches the right conclusion attached to the wrong or missing citation. |
Several of these score without a reference answer by using a model as a judge. A judge is cheap and repeatable, and it inherits its own bias, so a small human-labeled subset is kept to calibrate it. The same discipline applies to agents that plan and call tools, which are scored on task success, tool-call accuracy and trajectory quality instead — for more information visit the AI agent evaluation guide, or the AI testing framework for where these suites sit in a release process.
RAG Evaluation Frameworks: RAGAS, TruLens, DeepEval, ARES and Phoenix
The tooling is complementary rather than competing. A metrics library, a tracing layer and a continuous integration harness is the usual combination, and the choice matters far less than whether the suite runs automatically on every change.
| Project | What it is | Where it fits |
|---|---|---|
| RAGAS | Open-source metric library; the origin of the working definitions of faithfulness, answer relevance, context precision and context recall, several of which score without a reference answer. | A first read on retrieval and grounding quality, from a notebook or a script. |
| TruLens | Open-source evaluation and tracing library organized around the RAG triad of context relevance, groundedness and answer relevance. | Instrumenting a running application and watching per-trace scores. |
| DeepEval | A pytest-style harness with a broad metric catalog, model-as-judge scoring and red-teaming checks. | Turning the evaluation into a regression suite that runs on every change. |
| ARES | An automated evaluation approach (Saad-Falcon et al., NAACL 2024) that fine-tunes lightweight judges and corrects them with prediction-powered inference against a small human-labeled set. | Programs where the accuracy of the judge itself has to be defensible. |
| Arize Phoenix | Open-source tracing and observability with evaluations layered on top of captured spans. | Watching quality in production rather than only in a test harness. |
Blockify sits upstream of all of them: it changes what is in the index rather than how the index is scored, so any of these harnesses can measure the difference it makes. The retrieval stacks these evaluations run against are compared separately on the RAG frameworks page.
The Protocol Behind Iternal’s Measured Figures
Read against that scorecard, the benchmark on this page is a retrieval-side evaluation with a generation-side companion study.
- Retrieval. Mean cosine distance to the best-match result, 0.3624 for naive 1,000-character chunking against 0.1585 for distilled IdeaBlocks — a top-1 retrieval-quality measure on one corpus, with the embedding model, the vector database, the queries and the prompt held constant and only the representation changed. Distance rather than nDCG@10 because the comparison is like-for-like on a single enterprise corpus, not a leaderboard entry.
- Cost. Tokens per retrieved passage, roughly 303 for a naive chunk against roughly 98 for a distilled IdeaBlock, counted at a 4:1 character-to-token ratio. Accuracy and token cost are reported together because either one alone can be moved at the expense of the other.
- Generation. The medical-accuracy study scored combined accuracy and source fidelity across nine clinical questions — the same territory as faithfulness plus answer correctness — at a 261.11% average improvement and a 650% peak on diabetic-ketoacidosis management.
- Not reported. Recall@k, MRR and nDCG on a public BEIR split. These are controlled A/B comparisons of two representations of the same enterprise corpus, which is the question an enterprise buyer is asking, and a different question from how a retriever generalizes across public datasets.
Run the Same Comparison on Your Own Corpus
- Freeze every variable but one: same embedding model, same vector database, same top-k, same prompt. Change only how the documents are represented.
- Build a query set from real user questions. Fifty is enough to see a difference, and each one needs the passage that should have won written down.
- Score retrieval first, with recall@k and MRR or with best-match distance. A generation metric computed on broken retrieval measures the wrong thing.
- Score generation with faithfulness and answer relevance, then hand-check a sample so you know what the judge is rewarding.
- Record tokens per query beside every accuracy number. An accuracy gain bought with three times the context is a different result.
- Re-run the suite on every ingestion change and keep the history, so a regression is visible the day it lands rather than the quarter it is noticed.
Iternal runs this comparison on your own documents during a Blockify evaluation: one query set, scored against a chunked and a distilled version of the same corpus, reported on retrieval accuracy and tokens per query. The worked example is in the Blockify Performance Analysis, and Blockify is where an evaluation starts.
How to Cite This Page
Cite this data
Iternal Technologies. “Blockify Benchmark: RAG Data Optimization Performance.” iternal.ai/blockify-benchmarks. Retrieved 1970.
Data: Iternal Technologies, 1970. Free to cite with attribution.
Sources & References
- Iternal Technologies. Blockify® Performance Analysis Report for a Big Four Consulting Firm — the primary benchmark for vector accuracy (2.29X), token efficiency (3.09X), data reduction, and the 68.44X aggregate.
- Iternal Technologies. Evaluation of Blockify (Medical Accuracy Study) — the healthcare RAG evaluation across nine clinical questions (261.11% average, 650% peak).
- IDC. “Accelerating Efficiency and Driving Down IT Costs Using Data Duplication” — source for the 8:1–22:1 (average 15:1) enterprise data-duplication factor used in the enterprise-adjusted figures.
- Gartner. Lack of AI-Ready Data Puts AI Projects at Risk (2025) — context on why organizations abandon AI projects unsupported by AI-ready data.
- McKinsey. The economic potential of generative AI — context on the share of working hours generative AI can automate.
- Blockify open documentation (GitHub) — public reference for the ingestion, distillation, and governance pipeline.
- Blockify’s ingestion, distillation, and governance methods are patented by Iternal Technologies.
The complete methodology is published on this page and in the Blockify Performance Analysis — formulas, dataset scope, and derivations included.
Frequently Asked Questions
In documented benchmarks Blockify delivers up to 78X aggregate RAG accuracy improvement across enterprise deployments, driven by 2.29X more accurate vector search (a 56.26% precision gain) and up to 40X data reduction. On a single Big Four consulting-firm dataset the measured aggregate was 68.44X. All figures come from Iternal Technologies' Blockify Performance Analysis.
78X is the cross-deployment aggregate average. It combines the measured vector-search accuracy gain with data-volume reduction. On the transparent Big Four worked example, a 4.56X base improvement (2.29X vector accuracy × 2.00X word-count reduction) is multiplied by roughly 15X — IDC's average enterprise data-duplication factor (8:1 to 22:1) — for a 68.44X enterprise aggregate. The full calculation is published in the Blockify Performance Analysis.
Distilled IdeaBlocks average ~98 tokens versus ~303 tokens for a naive chunk, a 3.09X reduction in tokens processed per query. At 1 billion annual queries priced at $0.72 per million tokens (Llama 3.3 70B), that is an estimated $738,000 per year in savings, alongside lower latency and compute.
In an "Evaluation of Blockify" study across nine clinical questions, Blockify improved combined accuracy and source fidelity by 261.11% on average versus chunking, peaking at 650% on diabetic-ketoacidosis management and 500% on pneumonia lab tests. In one case, chunking recommended D5W dextrose as an initial IV fluid where isotonic saline is protocol; Blockify avoided the error.
On average Blockify shrinks a corpus to roughly 2.5% of its original size — up to 40X smaller. The measured base reduction on the Big Four dataset was 2.00X by word count, rising to 29.93X once enterprise-wide duplication (IDC's 15:1 average) is factored in. Smaller datasets mean fewer tokens, lower cost, and faster human review.
Yes. The full methodology, dataset scope, distance formulas, and token assumptions are published in the Blockify Performance Analysis and the healthcare evaluation, both linked on this page. Cite as: "Iternal Technologies, Blockify Benchmark, iternal.ai/blockify-benchmarks."
RAG evaluation measures a retrieval-augmented generation system in two halves. Retrieval is scored on whether the right passages came back, using recall@k, precision@k, MRR and nDCG, with BEIR (Thakur et al., NeurIPS 2021) as the standard portable benchmark of 18 public datasets across nine retrieval tasks. Generation is scored on whether the answer is supported by those passages, using faithfulness, answer relevance, context precision and context recall. A result is only readable when both halves are reported, alongside the tokens each query consumed.
On the retrieval side: recall@k (did the correct passage appear in the top k), precision@k (how much of what came back was relevant), MRR (how high the first correct passage ranked) and nDCG@k (position-discounted graded relevance). On the generation side: faithfulness (every claim traceable to a retrieved passage), answer relevance, context precision and context recall, whose working definitions were published with the RAGAS project. Token count per query belongs beside them, because an accuracy gain bought with a much larger context window is a different result.
They are complementary rather than competing. RAGAS covers retrieval and grounding metrics and scores several of them without a reference answer. DeepEval expresses the same checks as pytest-style tests that run in continuous integration. TruLens instruments a running application around context relevance, groundedness and answer relevance. ARES (Saad-Falcon et al., NAACL 2024) fine-tunes lightweight judges and corrects them against a small human-labeled set. Arize Phoenix adds tracing and evaluation in production. A metrics library plus a tracing layer plus a CI harness is the common combination.
Blockify changes the input to retrieval rather than the way retrieval is scored, so any evaluation harness can measure it. On the Big Four consulting-firm dataset the comparison held the embedding model, vector database, top-k and prompt constant and changed only the representation: naive 1,000-character chunks against distilled IdeaBlocks. Mean best-match cosine distance fell from 0.3624 to 0.1585, which is 2.29X more accurate vector search, while tokens per retrieved passage fell from roughly 303 to roughly 98.