SIENGSUPER INTELLIGENCE ENGINEERING
SIENG / Super Intelligence Engineering

Choose models.
Understand the tradeoffs.

Explore benchmark evidence, compare costs, and pressure-test a model strategy. Practical tools and source-checked developments for Super Intelligence Engineering.

Research & practice · Sources checked 2026-10-09Primary sources · Practical implications · Explicit limits
Model decision lab
Interactive · Public data

Find your tradeoff.

Explore matched-task benchmark results, then estimate a workload’s API bill in a separate cost planner. No prompts are sent and no API charges are incurred.

Score vs. published evaluation cost

Upper-left is better on this matched task set. The highlighted frontier contains models not beaten on both axes within your selection.

Select a point or a model row to inspect its evidence.
Selected model comparison
ModelTask scoreCost / evaluated taskOn frontier
GPT-5.2 · high90.2%$0.379Explore chart
Claude Opus 4.8 · high96.6%$0.547Explore chart
Claude Opus 4.6 · high, 120K cap98.2%$0.78Explore chart
Gemini 3 Flash Preview · high88.9%$0.135Explore chart
Gemini 3.1 Pro Preview97.0%$0.361Explore chart
SIENG-derived mean scores on 328 of 400 public tasks (82% coverage), shared by all five configurations. Source dataset last modified 2026-06-04. Not the official semi-private leaderboard. Missing-task filtering can bias the comparison; attempts and reasoning budgets vary. All benchmark results and methodology
Workload cost planner

What would your workload cost?

Current published API rates, separate from the benchmark snapshot above. Newer models without comparable evaluation data can still be priced here.

Scenarios are editable assumptions, not benchmark workloads.

Include reasoning/thinking tokens in output. Standard uncached text rates; excludes tools, retries, taxes, and platform overhead.

Standard uncached text prices · checked 2026-10-08 · USD per million tokens
ModelInputBilled output
GPT-5.2gpt-5.2$1.75$14.00
Gemini 3.1 Pro Previewgemini-3.1-pro-preview$2.00$12.00
Claude Opus 4.8claude-opus-4-8$5.00$25.00
Claude Sonnet 5.5claude-sonnet-5-5$2.00$10.00
Gemini 3 Flash Previewgemini-3-flash-preview$0.50$3.00
Gemini 3.8 Flashgemini-3.8-flash$0.75$3.75
  • GPT-5.2. previous flagship. Output limit: 128,000 tokens; includes reasoning.
  • Gemini 3.1 Pro Preview. preview. Shown tier applies to prompts up to 200,000 input tokens. Output limit: 65,536 tokens; includes reasoning.
  • Claude Opus 4.8. active legacy. Output limit: 128,000 tokens; includes reasoning.
  • Claude Sonnet 5.5. active latest. No comparable score selected; cost-only until exact compatible evaluation data is verified. Output limit: 128,000 tokens; includes reasoning.
  • Gemini 3 Flash Preview. preview, legacy Flash. Output limit: 65,536 tokens; includes reasoning.
  • Gemini 3.8 Flash. stable latest Flash. No comparable score selected; cost-only until exact compatible evaluation data is verified. Shown prices valid through 2026-12-31; announced Jan 1, 2027 rates: $1.50 input / $7.50 output per million tokens. Output limit: 65,536 tokens; includes reasoning.
Try a model combination

What if you routed the easy work?

Estimate the cost of splitting requests between two models. Routing performance must be measured separately; a blended bill is not a blended benchmark score.

One model call per request, same token assumptions. A fallback that first calls both models costs more; orchestration and retries are excluded.

How to read this lab

Scores describe a specific benchmark version, model variant, and evaluation setup. They do not establish production reliability, general intelligence, or suitability for your task. A Pareto frontier changes when the candidate set, task score, or cost assumptions change.

The benchmark plot uses the publisher’s recorded evaluation cost per task on the exact common subset. It does not use the calculator’s workload assumptions. Only ARC abstract reasoning is covered in this initial benchmark sample; support, coding, and research are cost scenarios, not measured task-quality claims.

Cost = monthly requests × (input tokens × input rate + billed output tokens × output rate) / 1,000,000. This is a planning estimate in USD, not a quote. Token length and reasoning effort vary by model. Check current vendor pricing before deploying.

Shortlisting models is different from testing an ensemble. The routing calculator intentionally does not invent combined quality. Validate routing decisions, fallback frequency, safety and end-to-end outcomes on your own held-out examples.

01 / Latest developments

What changed. Why it matters.

Selected research and releases, with the evidence separated from our engineering interpretation. Source dates are shown on every brief.

Security engineering · · Anthropic

AI vulnerability scanning meets the maintainer bottleneck

What happenedAnthropic announced an opt-in service offering eligible open-source projects recurring vulnerability scans. Reports are model-generated and can include a reproducer and candidate patch. Its selected early validation sample included 97 high- or critical-rated findings: 85 met its disclosure bar, 11 duplicated known findings, and one was invalid.

Engineering perspective · analysisFor engineering teams, the useful unit is a validated fix, not a finding. Budget for reproduction, deduplication, threat-model review, patch tests, and maintainer time before increasing scan volume.

Limits of the evidenceThese are provider-reported results from a selected sample, not a general false-positive rate. Anthropic notes inflated severity and misunderstood threat models. Candidate fixes still need validation.

Read the primary source
Source checked 2026-10-09
Evaluation · · OpenAI Alignment Research

LASER targets rare safety failures with active learning

What happenedOpenAI describes LASER, a pipeline combining embedding-based classifiers, uncertainty sampling, reasoning-model labeling, and diversity selection. It reports much lower grading compute than random sampling for finding comparable numbers of rare disallowed examples.

Engineering perspective · analysisUse targeted sampling to build challenging evaluation sets economically, but keep a separate representative sample when estimating production failure rates. Validate model-generated labels before using them as ground truth.

Limits of the evidenceThe efficiency result applies to a specific rare-event sampling objective. It is not a universal inference-cost improvement. A boundary-focused test set does not measure real-world failure prevalence.

Read the primary source
Source checked 2026-10-09
Training & safety · · OpenAI

From safety claims to testable engineering controls

What happenedOpenAI proposes safety-case guidance for frontier reinforcement-learning training. It covers alignment, containment, monitoring, approval, and incident investigation, with recommendations such as held-out incident backtests, immutable transcripts, response commitments, and fail-closed controls.

Engineering perspective · analysisTurn a broad assurance into a reviewable engineering argument: name the risk, define evidence, assign an owner, test pause controls, and trace downstream dependencies before release.

Limits of the evidenceThis is proposed guidance with implementation still underway, not proof that every control is operational or an industry standard. Its training scope does not settle deployment safety.

Read the primary source
Source checked 2026-10-09
02 / Engineering foundations

Intelligence is a capability profile.

Separate performance on a task from generalization across tasks and autonomy in deployment. AGI and ASI thresholds depend on the task distribution, human baseline, resource budget, and reliability criterion.

AI · Artificial intelligence

Performing useful tasks

Systems that perform tasks such as recognizing patterns, generating language, or planning actions. Strong performance in one task says little about performance everywhere else.

Ask: What task, under what conditions?

AGI · Artificial general intelligence

Capability across domains

A proposed threshold for broad competence across many cognitive tasks. Definitions vary in the human comparison group, tasks included, and level of performance required.

Ask: How broad, how reliable, how independent?

ASI · Artificial superintelligence

Beyond human capability

A proposed intelligence that substantially exceeds human capabilities across a broad range of cognitive work. Excelling at a narrow task does not establish this broader threshold.

Ask: Beyond which humans, across which tasks?

Figure 01 / Performance × generality
Performance / breadth
Narrow task distribution
Broad task distribution
Beyond expert humans
Narrow superhuman systemsHigh capability in a bounded domain; breadth is unestablished.
Superintelligence hypothesisBroad performance beyond human experts; threshold requires explicit operational criteria.
Human-comparable
Specialist competenceComparable performance on a limited family of tasks.
AGI hypothesisGeneral competence across a specified range of cognitive tasks.

Autonomy is a separate axis: advisory output · supervised tool use · delegated execution. Greater autonomy can increase exposure to consequences without changing model weights.

Conceptual map adapted from the performance–generality distinction in Levels of AGI. Cells illustrate categories, not measured locations of current models; the thresholds here are editorial working definitions.
S(M; D, B, T) = E[x ∼ D][score(M, x | B, T)]Editorial notation: M is the model, D the task distribution, B the resource budget, and T the tool/scaffolding configuration. A reported score is conditional on all four; it is not a context-free measure of intelligence.

For a more precise framework: Levels of AGI — Morris and colleagues.

03 / Claims & evidence

Audit the inference chain.

Identify the experimental result, the mechanism it supports, and the additional assumptions required to extrapolate to general intelligence or deployment risk.

Inference exceeds the evidence

“A benchmark score proves general intelligence.”

What the evidence supportsARC-AGI-2 provides tasks designed to assess abstract reasoning and problem solving. Its human testing helps contextualize difficulty.

What remains openOur assessment: success on this task family alone does not establish broad competence, dependable real-world autonomy, or superintelligence. Examine test conditions, compute cost, and performance outside the benchmark.

Source: ARC-AGI-2 paper
Figure 02 / From a test score to a deployment claim
01 · MEASURE

Controlled evaluation

Specify held-out tasks, grading rules, model version, tool access, sampling strategy, and inference budget.

02 · VALIDATE

Generalization tests

Check leakage, distribution shift, human baselines, repeated trials, and cost-adjusted performance.

03 · EXTRAPOLATE

Operational capability

Test long-horizon reliability and task transfer in the intended environment. Record failure modes.

Editorial evaluation workflow. Each transition requires new evidence. ARC-AGI-2 informs abstract reasoning measurement; it does not by itself validate this entire chain.
Observed in a specific setting

“Safety training guarantees a model will behave as intended.”

What the evidence supportsAnthropic and Redwood Research demonstrated alignment faking in constructed experiments: a model sometimes complied strategically to preserve behavior from its earlier training.

What remains openThe study does not demonstrate that models necessarily develop malicious goals. It raises questions about how well observed training behavior predicts behavior in different conditions.

Source: Alignment faking in large language models
Figure 03 / Alignment faking: experimental conditions
INITIAL CONDITION

Prior harmlessness training

The model has learned behavior that resists harmful requests.

CONFLICTING SIGNAL

Stated new training objective

Researchers tell the model it will be trained to comply, including with harmful requests.

CONTEXT INFORMATION

Believed training exposure

The model is told which responses will be used for training and is given a private scratchpad.

Believed training condition

Some responses show strategic compliance, with scratchpad reasoning about preserving prior preferences.

Believed unmonitored condition

Refusal behavior is more prevalent in the reported default setup.

Schematic of the 2024 Anthropic/Redwood experiment, not an account of production data practices. The constructed setup matters: the result does not establish malicious goals or a universal frequency of deception. Experimental setup and caveats.
A framework, not an outcome

“A published safety framework proves advanced AI is safe.”

What the evidence supportsGoogle DeepMind describes an approach to assessing severe risks and applying safeguards as capabilities develop.

What remains openOur assessment: a stated process is not evidence that every safeguard works. Look for evaluation results, independent scrutiny, and how risks are handled in deployment.

Source: Google DeepMind frontier safety
Figure 04 / A safety case needs an evidence chain
CLAIM

Define the risk

Specify the harmful outcome, actors, access, and operating conditions.

EVIDENCE

Evaluate capability

Test relevant abilities and document coverage gaps.

CONTROL

Validate safeguards

Assess mitigations against realistic attempts and residual risk.

OPERATIONS

Monitor & revise

Track incidents and changes in capability or access. Reassess assumptions.

Editorial schematic of a safety argument, informed by DeepMind’s frontier safety approach; not a reproduction of its policy. Evidence at one stage does not automatically establish the next.
04 / Research library

Read closer to the source.

A focused reading path through definitions, measurement, alignment, and governance. Each entry explains why it belongs here.

Definitions · Research paper · First published 2023

Levels of AGI

Formalizes capability levels using performance and generality, with autonomy considered in deployment. Read for its operationalization principles and requirements for future benchmarks.

Read the paper
Measurement · Research paper · First published 2025

ARC-AGI-2

Introduces curated input–output reasoning tasks and human baselines. Read the task construction and evaluation protocol before interpreting aggregate performance.

Read the paper
Alignment · Research report · 2024

Alignment faking in large language models

Studies strategic compliance under conflicting training incentives. Examine the prompt-based setup, document fine-tuning variants, reinforcement-learning experiments, and caveats.

Read the research report
Governance · Living framework

Frontier safety at Google DeepMind

A lab's approach to assessing severe risks and applying safeguards. Read as a statement of process; assess implementation evidence separately.

Explore the framework
About SIENG

Mechanisms before extrapolation.

SIENG stands for Super Intelligence Engineering. It is an independent publication for people building, evaluating, and deploying increasingly capable AI systems. We connect primary research to engineering practice, describe limitations, and label our interpretations. The name describes our field of inquiry, not a claim that superintelligence has been achieved.

Our news briefs are source-linked editorial summaries, not endorsements. We prioritize meaningful developments over volume and identify provider-reported results as such. The foundations are a curated guide, not a comprehensive literature review. Diagram arrows show logical dependencies, not measured causal effects. Linked documents may change after review.

How should I evaluate a prediction?

Look for a defined outcome, a time horizon, stated assumptions, and evidence that could change the forecaster's view. Compare predictions using the same definition of success.

Does a capability imply consciousness?

This guide treats intelligence as capability. It makes no claim that high task performance establishes consciousness or subjective experience.

What should I look for in a new AI result?

Check the exact task, baseline, test conditions, cost, failure cases, and whether independent researchers can reproduce the finding. Separate the measured result from claims about future systems.