ZipMind

Towards the Silicon Brain

PEREA: a Neuro-symbolic Foundation

Two core technologies set PEREA apart: a Human brain-inspired neuro-symbolic architectural design & Focalized Model Compression.

Human brain-inspired neuro-symbolic architectural design

  • More robust visual understanding
  • Much stronger logical reasoning
Neuro-symbolic architecture: symbolic circuit hemisphere and neural hemisphere
Focalized model compression funnel

Focalized Model Compression

  • Extreme compression rate
  • Trivial performance drop

How good are we?

Offering superior visual reasoning performance.

Chart Understanding (CharXiv) results with PEREA-1.0 near the top of frontier models
Chart Understanding (CharXiv) — PEREA-1.0 scores 91.2, standing shoulder-to-shoulder with the strongest frontier models.
Intelligence Test (VisuLogic) results with PEREA-1.0 ranked first above the human baseline
Intelligence Test (VisuLogic) — PEREA-1.0 leads at 52.8, surpassing the human baseline (51.4).

Not just performant, but also cost-efficient.

Cost Efficiency by Model — PEREA delivers far more accuracy per dollar than frontier models

What we offer?

From an instant web chat to production APIs and bespoke compression — put PEREA to work your way.

Document Analyst

Leverage PEREA's powerful visual reasoning to intelligently understand and analyze documents packed with complex charts and tables.

Try it now →

PEREA API

Bring PEREA's visual reasoning into your own product with a simple, scalable API built for developers.

Try it now →

Compression API

We focalize and compress your SoTA models for target hardware — extreme compression rate with trivial performance drop.

Try it now →
Research

ZipMind Research

Introductions to the technology and products behind the Silicon Brain.

← Research

Introducing PEREA-1.0

A neuro-symbolic foundation model for visual reasoning

Towards the Silicon Brain

We officially announce the release of PEREA-1.0, the world’s first human brain–inspired neuro-symbolic foundation model — or, as we like to call it, the Silicon Brain. Thanks to its novel neuro-symbolic architecture, PEREA delivers stunning visual reasoning: it surpasses humans on an intelligence test, achieves top-3 results on chart understanding, and does so with superior cost efficiency.

Chart understanding that matches the frontier

Scientific and analytical charts are a stress test for visual models: axes, legends, nested panels, and questions that need more than OCR.

On CharXiv, PEREA-1.0 scores 91.2, sitting with the strongest frontier systems — just behind Claude Mythos Preview (93.2), essentially tied with Kimi K3 (91.3), and ahead of Claude Opus 4.8, GPT-5.x, Gemini, and Qwen variants in this comparison.

Chart Understanding (CharXiv) results
Chart Understanding (CharXiv) — PEREA-1.0 overall score 91.2.

Harder than charts: visual intelligence

We further challenge ourselves with VisuLogic, an intelligence test that requires strong abstract visual reasoning.

Here PEREA-1.0 leads at 52.8, becoming the first model to officially surpass the strong human baseline of 51.4 (see the official leaderboard), and sitting well clear of other AI counterparts.

Intelligence Test (VisuLogic) results
Intelligence Test (VisuLogic) — PEREA-1.0 at 52.8, above human (51.4).

For us, this is the sharper claim: not only competitive on chart QA, but first on a benchmark where humans are still the bar.

Not just performant — also cost-efficient

Frontier accuracy alone is not enough if every query burns the budget. We measure cost efficiency as CharXiv accuracy percentage points per full-dataset USD — higher is better.

PEREA scores 13.20, far ahead of GPT-5.6 (3.41), Claude Fable (2.50), Seed 2.1 Pro (1.84), and Kimi K3 (1.46). Roughly ~4× GPT-5.6 and ~9× Kimi K3 on this metric.

Cost Efficiency by Model
Cost efficiency — PEREA delivers more accuracy per dollar.

That is the product shape we want: SoTA-level visual reasoning without SoTA-level spend.

What makes PEREA different

Two ideas sit under the numbers:

  1. Human brain–inspired neuro-symbolic architectural design — offering much stronger visual understanding and logical reasoning, closer to its carbon-based ancestor.
  2. Focalized model compression (Foka) — extreme compression with trivial performance drop, so the same capability can land on the hardware you actually have.

Together they push toward what we call the Silicon Brain: intelligence that is both sharp and efficient.

What’s next

PEREA-1.0 is in private preview. If you are building document analysis, chart-heavy workflows, or need visual reasoning behind an API, we would like to hear from you.

Request access: zipmind2026@gmail.com

Towards the Silicon Brain.

← Research

Introducing Foka

Focalized model compression for faster, smaller, and more energy-efficient AI

Foka

We are introducing Foka, ZipMind's focalized model compression method for turning large models into practical deployment gains.

Foka is designed around the outcomes that matter after a model leaves the laboratory: lower energy consumption, a smaller memory footprint, and higher inference speed — without giving up more model quality than necessary. Across BERT, RoBERTa, VGG19, Informer, and Llama-3-8B, Foka consistently reaches the most efficient operating point in our comparison.

  • Up to 98% lower energy consumption
  • Up to 81% lower memory usage
  • Llama-3-8B speed-up at perplexity 16.47

Compression that translates into real savings

Model compression is often reported as a parameter count or sparsity ratio. Those numbers matter, but they do not fully describe whether a model is actually cheaper to run or easier to deploy.

We therefore evaluate Foka directly through energy consumption and memory footprint, comparing it with the original base model and a matched-performance state-of-the-art model compression reference.

Energy consumption and memory footprint comparison
Energy and memory comparison across specialised models and Llama-3-8B. Lower is better.

Across the specialised models, Foka reduces energy consumption to only a small fraction of the uncompressed baseline: 2% for BERT, 2% for RoBERTa, 4% for VGG19, and 3% for Informer. The matched SOTA configurations consume between 10% and 17% of baseline energy.

The same pattern appears in memory usage. Foka reduces BERT from 420 MB to 80 MB, RoBERTa from 500 MB to 95 MB, VGG19 from 548 MB to 135 MB, and Informer from 180 MB to 40 MB. In every case, Foka also uses substantially less memory than the matched SOTA comparison.

The gains continue at large-model scale. For Llama-3-8B, Foka uses 33% of baseline energy, versus 50% for the matched SOTA method. Its memory footprint falls from 15,360 MB to 5,120 MB — one-third of the original model and 2,560 MB below the matched SOTA result.

This is the kind of compression we care about: not sparsity for its own sake, but savings that directly affect infrastructure cost, hardware fit, and deployment efficiency.

Three times faster Llama-3-8B, with quality still under control

Efficiency is only valuable when the model remains useful. We therefore measure how perplexity changes as Llama-3-8B is accelerated from its original operating point to 2× and 3× speed-up.

Perplexity versus speed-up comparison for Llama-3-8B
Perplexity versus speed-up on Llama-3-8B. Lower perplexity is better; the dashed line marks the quality limit of 25.

The uncompressed model starts at a perplexity of 5.72. At 2× speed-up, Foka reaches 6.27 — an increase of only 0.55 — showing that substantial acceleration is possible with a relatively small quality shift.

At the more demanding 3× speed-up point, the differences become much clearer. Foka reaches a perplexity of 16.47, compared with 24.93 for Wanda++, 26.47 for SparseGPT, 29.66 for Thanos, 42.40 for FLAP, and 50.39 for SlimGPT.

Foka is the strongest result in the comparison and remains comfortably below the perplexity limit of 25. At 3× speed-up, it is 8.46 perplexity points better than Wanda++ and 10.00 points better than SparseGPT.

One method, multiple model families

Foka is not evaluated on a single architecture or task family. Our results span language encoders, computer vision, time-series forecasting, and large language models.

  1. Language encoders — BERT and RoBERTa demonstrate that the efficiency gains extend across transformer-based language understanding models.
  2. Computer vision — VGG19 shows that Foka can reduce the resource footprint of convolutional visual models.
  3. Time-series forecasting — Informer extends the comparison beyond conventional language and vision workloads.
  4. Large language models — Llama-3-8B tests Foka at foundation-model scale under aggressive speed-up targets.

This breadth matters because redundancy does not appear in exactly the same form across every model. A useful model compression method must remain effective across different structures, workloads, and deployment constraints.

What makes Foka different

Foka is built around a deployment-first view of model compression.

  1. Compression should be focalized — Not every part of a model contributes equally to its final capability. Foka concentrates compression where redundancy can be removed most safely, rather than treating the model as uniformly expendable.
  2. The operating point matters more than the sparsity number — The goal is not simply to remove as many parameters as possible. It is to find a better balance between speed, memory, energy, and quality.
  3. Efficiency should survive aggressive settings — Many methods look competitive at moderate compression but deteriorate sharply as the target speed-up increases. Foka's advantage is most visible at 3×, where it maintains substantially lower perplexity than the alternatives.

Together, these principles make Foka a practical compression layer for models that need to move from large-scale experimentation to cost-efficient deployment.

From large models to deployable intelligence

Modern AI systems continue to become more capable — but also larger, slower, and more expensive to operate.

For many applications, the challenge is no longer whether a model can perform the task. It is whether the same capability can fit within the available GPU memory, latency target, energy budget, or deployment hardware.

Foka addresses this gap by making efficiency a first-class model capability. On specialised models, it reduces resource requirements to a fraction of the original system. On Llama-3-8B, it enables significantly more aggressive acceleration while preserving substantially better perplexity than existing model compression methods.

That is the product shape we want: models that retain their intelligence while becoming practical enough to deploy.

What's next

We are extending Foka to larger foundation models, additional hardware targets, and more end-to-end deployment settings. Our next evaluations will focus on measured latency, energy, throughput, and model quality across a broader range of production workloads.

Foka is also a key part of ZipMind's broader goal: building intelligence that is not only capable, but efficient enough to run on the hardware people actually have.

If you are deploying models that are too large, too slow, or too expensive for their target environment, we would like to hear from you.

Request access: zipmind2026@gmail.com

Towards the Silicon Brain.