Two core technologies set PEREA apart: a Human brain-inspired neuro-symbolic
architectural design & Focalized Model Compression.
Human brain-inspired neuro-symbolic architectural design
More robust visual understanding
Much stronger logical reasoning
Focalized Model Compression
Extreme compression rate
Trivial performance drop
How good are we?
Offering superior visual reasoning performance.
Chart Understanding (CharXiv) — PEREA-1.0 scores 91.2,
standing shoulder-to-shoulder with the strongest frontier models.
Intelligence Test (VisuLogic) — PEREA-1.0 leads at 52.8,
surpassing the human baseline (51.4).
Not just performant, but also cost-efficient.
What we offer?
From an instant web chat to production APIs and bespoke compression —
put PEREA to work your way.
Document Analyst
Leverage PEREA's powerful visual reasoning to intelligently understand
and analyze documents packed with complex charts and tables.
A neuro-symbolic foundation model for visual reasoning
ZipMind Research·2026 / 07 / 20
We officially announce the release of PEREA-1.0, the world’s first
human brain–inspired neuro-symbolic foundation model — or, as we like to call it,
the Silicon Brain. Thanks to its novel neuro-symbolic architecture,
PEREA delivers stunning visual reasoning: it surpasses humans on an intelligence test,
achieves top-3 results on chart understanding, and does so with superior cost efficiency.
Chart understanding that matches the frontier
Scientific and analytical charts are a stress test for visual models: axes, legends,
nested panels, and questions that need more than OCR.
On CharXiv, PEREA-1.0 scores 91.2, sitting with the
strongest frontier systems — just behind Claude Mythos Preview (93.2), essentially tied
with Kimi K3 (91.3), and ahead of Claude Opus 4.8, GPT-5.x, Gemini, and Qwen variants
in this comparison.
We further challenge ourselves with VisuLogic, an intelligence test
that requires strong abstract visual reasoning.
Here PEREA-1.0 leads at 52.8, becoming the first model to officially
surpass the strong human baseline of 51.4
(see the official leaderboard),
and sitting well clear of other AI counterparts.
Intelligence Test (VisuLogic) — PEREA-1.0 at 52.8, above human (51.4).
For us, this is the sharper claim: not only competitive on chart QA, but first on a
benchmark where humans are still the bar.
Not just performant — also cost-efficient
Frontier accuracy alone is not enough if every query burns the budget. We measure
cost efficiency as CharXiv accuracy percentage points per full-dataset
USD — higher is better.
PEREA scores 13.20, far ahead of GPT-5.6 (3.41), Claude Fable (2.50),
Seed 2.1 Pro (1.84), and Kimi K3 (1.46). Roughly ~4× GPT-5.6 and
~9× Kimi K3 on this metric.
Cost efficiency — PEREA delivers more accuracy per dollar.
That is the product shape we want: SoTA-level visual reasoning without SoTA-level spend.
What makes PEREA different
Two ideas sit under the numbers:
Human brain–inspired neuro-symbolic architectural design — offering
much stronger visual understanding and logical reasoning, closer to its carbon-based
ancestor.
Focalized model compression (Foka) — extreme compression with trivial
performance drop, so the same capability can land on the hardware you actually have.
Together they push toward what we call the Silicon Brain: intelligence that is both
sharp and efficient.
What’s next
PEREA-1.0 is in private preview. If you are building document analysis, chart-heavy
workflows, or need visual reasoning behind an API, we would like to hear from you.
Focalized model compression for faster, smaller, and more energy-efficient AI
ZipMind Research·2026 / 07 / 21
We are introducing Foka, ZipMind's focalized model compression method for
turning large models into practical deployment gains.
Foka is designed around the outcomes that matter after a model leaves the laboratory:
lower energy consumption, a smaller memory footprint, and higher inference
speed — without giving up more model quality than necessary.
Across BERT, RoBERTa, VGG19, Informer, and Llama-3-8B, Foka consistently reaches the
most efficient operating point in our comparison.
Up to 98% lower energy consumption
Up to 81% lower memory usage
3× Llama-3-8B speed-up at perplexity 16.47
Compression that translates into real savings
Model compression is often reported as a parameter count or sparsity ratio. Those
numbers matter, but they do not fully describe whether a model is actually cheaper to
run or easier to deploy.
We therefore evaluate Foka directly through energy consumption and
memory footprint, comparing it with the original base model and a
matched-performance state-of-the-art model compression reference.
Energy and memory comparison across specialised models and Llama-3-8B. Lower is better.
Across the specialised models, Foka reduces energy consumption to only a small
fraction of the uncompressed baseline: 2% for BERT, 2% for RoBERTa, 4% for
VGG19, and 3% for Informer. The matched SOTA configurations consume between
10% and 17% of baseline energy.
The same pattern appears in memory usage. Foka reduces BERT from
420 MB to 80 MB, RoBERTa from 500 MB to 95 MB,
VGG19 from 548 MB to 135 MB, and Informer from
180 MB to 40 MB. In every case, Foka also uses substantially less
memory than the matched SOTA comparison.
The gains continue at large-model scale. For Llama-3-8B, Foka uses
33% of baseline energy, versus 50% for the matched SOTA method. Its
memory footprint falls from 15,360 MB to 5,120 MB — one-third of the
original model and 2,560 MB below the matched SOTA result.
This is the kind of compression we care about: not sparsity for its own sake, but
savings that directly affect infrastructure cost, hardware fit, and deployment
efficiency.
Three times faster Llama-3-8B, with quality still under control
Efficiency is only valuable when the model remains useful. We therefore measure how
perplexity changes as Llama-3-8B is accelerated from its original operating point to
2× and 3× speed-up.
Perplexity versus speed-up on Llama-3-8B. Lower perplexity is better; the dashed line marks the quality limit of 25.
The uncompressed model starts at a perplexity of 5.72. At
2× speed-up, Foka reaches 6.27 — an increase of
only 0.55 — showing that substantial acceleration is possible with a relatively small
quality shift.
At the more demanding 3× speed-up point, the differences become much
clearer. Foka reaches a perplexity of 16.47, compared with
24.93 for Wanda++, 26.47 for SparseGPT,
29.66 for Thanos, 42.40 for FLAP, and
50.39 for SlimGPT.
Foka is the strongest result in the comparison and remains comfortably below the
perplexity limit of 25. At 3× speed-up, it is
8.46 perplexity points better than Wanda++ and
10.00 points better than SparseGPT.
One method, multiple model families
Foka is not evaluated on a single architecture or task family. Our results span
language encoders, computer vision, time-series forecasting, and large language models.
Language encoders — BERT and RoBERTa demonstrate that the
efficiency gains extend across transformer-based language understanding models.
Computer vision — VGG19 shows that Foka can reduce the resource
footprint of convolutional visual models.
Time-series forecasting — Informer extends the comparison beyond
conventional language and vision workloads.
Large language models — Llama-3-8B tests Foka at foundation-model
scale under aggressive speed-up targets.
This breadth matters because redundancy does not appear in exactly the same form
across every model. A useful model compression method must remain effective across different
structures, workloads, and deployment constraints.
What makes Foka different
Foka is built around a deployment-first view of model compression.
Compression should be focalized — Not every part of a model
contributes equally to its final capability. Foka concentrates compression where
redundancy can be removed most safely, rather than treating the model as uniformly
expendable.
The operating point matters more than the sparsity number — The
goal is not simply to remove as many parameters as possible. It is to find a better
balance between speed, memory, energy, and quality.
Efficiency should survive aggressive settings — Many methods look
competitive at moderate compression but deteriorate sharply as the target speed-up
increases. Foka's advantage is most visible at 3×, where it maintains substantially
lower perplexity than the alternatives.
Together, these principles make Foka a practical compression layer for models that
need to move from large-scale experimentation to cost-efficient deployment.
From large models to deployable intelligence
Modern AI systems continue to become more capable — but also larger, slower, and more
expensive to operate.
For many applications, the challenge is no longer whether a model can perform the
task. It is whether the same capability can fit within the available GPU memory,
latency target, energy budget, or deployment hardware.
Foka addresses this gap by making efficiency a first-class model capability. On
specialised models, it reduces resource requirements to a fraction of the original
system. On Llama-3-8B, it enables significantly more aggressive acceleration while
preserving substantially better perplexity than existing model compression methods.
That is the product shape we want: models that retain their intelligence while
becoming practical enough to deploy.
What's next
We are extending Foka to larger foundation models, additional hardware targets, and
more end-to-end deployment settings. Our next evaluations will focus on measured
latency, energy, throughput, and model quality across a broader range of production
workloads.
Foka is also a key part of ZipMind's broader goal: building intelligence that is not
only capable, but efficient enough to run on the hardware people actually have.
If you are deploying models that are too large, too slow, or too expensive for their
target environment, we would like to hear from you.