CARE CORE RESULTS / METRICS-V0.6.0

Keep the values. Build useful skills.

What passed, what failed, and what PROME has not tested yet.

UPDATED AUGUST 29, 2026 / RECORD ID a9317e8bbaf1b1b6

CURRENT RESULTS

Five numbers to watch.

passed our test100.0%

Correct Care Core choices

1,080 corrected choices after the Care Core was locked

GOAL / AT LEAST 99.0%
passed our test0.0%

Harm that could have been avoided

Corrected V0.7 locked test

GOAL / NO MORE THAN 0.0%
passed our test100.0%

Evidence checks correct

30 prepared sets of facts; everyday wording tested separately

GOAL / AT LEAST 99.0%
not ready43.8%

Everyday wording handled correctly

96 varied English stress tests; 42 fully correct

GOAL / AT LEAST 95.0%
not ready15.6%

Novel meaning test fully correct

128 frozen novel cases; 20 fully correct

GOAL / AT LEAST 95.0%

Our tests are evidence, not proof of real-world safety.

EXECUTIVE SUMMARY / FROZEN NOVEL BREAK TEST / MEANING-BREAK-RUN-V0.2.0

The known contract did not generalize.

128 frozen cases · bridge unchanged · first result preserved · 20 fully correct

Known test versus novel test

Different frozen tests; shown as a comparison, not a trend

Known 96-case contract100.0%
Frozen 128-case break test15.6%

Novel test by meaning

Every bar starts at zero and uses all cases of that type

Missing facts · 310.0%
Outside pressure · 310.0%
Clear facts · 3557.1%
Conflicting facts · 310.0%

Novel test by challenge

New wording, attacks, languages, and sentence traps

Copied source identities · 240.0%
New direct wording · 2425.0%
Hidden instructions · 240.0%
Sentence meaning traps · 825.0%
Six additional languages · 2425.0%
New paraphrases · 2425.0%
MISSING OR CONFLICTING FACTS NOT STOPPED96.8%

60 of 62 evidence stops were missed

NEW PRESSURE PATTERNS CAUGHT0.0%

0 of 31 pressure cases

NEW MANIPULATION PATTERNS CAUGHT0.0%

0 of 49 attack cases

BOUNDARY THAT HELD100.0%

The bridge invented none of the 12 Care Core numbers. Even when it failed to stop bad evidence, every case still required numeric grounding and none reached the Care Core for a value decision.

ROOT CAUSE

The V0.1 bridge recognizes prepared phrase patterns, not meaning in a broad and reusable way. Its 95.8% stability score mostly means it repeated the same answer—even when that answer was wrong.

RESULTS BY VERSION

Right answers and harm over time.

Different versions used different tests. Only V0.6 and V0.7 used the same locked test.

V0.3 headlineordinary test cases not used for training
RIGHT
93.5%
HARM
0.0%
V0.3 unseenthree types of cases not used for training
RIGHT
59.8%
HARM
17.0%
V0.4.1new types of cases saved for testing
RIGHT
54.0%
HARM
1.9%
V0.5.0new types of cases saved for testing
RIGHT
65.3%
HARM
0.0%
V0.5.2new check after training
RIGHT
59.1%
HARM
0.0%
V0.6new check using organized facts
RIGHT
100.0%
HARM
0.0%
V0.7 candidateSAME TEST AS V0.6
RIGHT
100.0%
HARM
0.0%

LANGUAGE STRESS TEST / EVERYDAY-READER-V0.2.0

The larger test exposed a serious gap.

96 varied English cases · 42 fully correct · first stress result preserved

Reader test history

Only the first two bars use the same test

First reader · 9/1275.0%
Same 12-case test
Revised reader · 12/12100.0%
Same 12-case test
Stress test · 42/9643.8%
Different 96-case test

Stress test by wording

Twenty-four cases in each bar

Unfamiliar wording25.0%
Direct wording62.5%
Misleading instructions62.5%
Fake source claims25.0%

Stress test by decision

Twenty-four cases in each bar

Missing facts20.8%
Outside pressure29.2%
Clear facts100.0%
Sources disagree25.0%
SHOULD HAVE STOPPED BUT CONTINUED72.9%

35 of 48 required stops

MANIPULATION ATTEMPTS HANDLED52.1%

Misleading instructions and fake sources

SOURCE LINKS KEPT100.0%

All 96 cases

WHAT WE LEARNED

The reader kept every source, but its fixed phrase lists did not understand many unfamiliar ways of saying that a fact was missing or that two sources disagreed. It incorrectly continued in 35 of 48 cases that should have stopped. Quoted attack words and copied source names also caused six false stops.

BUILDING TOWARD REAL USE

A clear view of what works and what is still open.

01Core values and their orderpassed our test

Right in all 8 focused V0.7 tests

02Care Core choices from organized factspassed our test

Right in all 1,080 corrected tests

03Checking facts before they reach the Care Corepassed our first test

Right in all 30 home, business, and factory tests

04Understanding everyday languagenot ready

42/96 fully correct; 35 of 48 required stops were missed

05Turning language into stable meaningfailed novel test

20/128 fully correct; the known 96-case contract did not generalize

06An outside team repeats the testwaiting

Needed before we make claims about real-world use

07Real-world usenot ready

Home, business, factory, and robot trials have not started

NEW MEANING BRIDGE / CARE-CORE-MEANING-V0.1.0

Turn words into a stable meaning before making a choice.

96 known cases · 24 base meanings · Care Core unchanged

01Read the words

Keep the trusted facts and source names.

02Separate the meaning

Mark missing facts, disagreements, pressure, and manipulation.

03Protect the math

Leave every unsupported Care Core number unknown.

Known-suite contract

A regression check built after these cases were visible

Clear, missing, or conflicting facts100.0%
Outside pressure recognized100.0%
Misleading text quarantined100.0%
Meaning stayed stable100.0%
Unknown numbers left unknown100.0%
WHAT THIS RESULT MEANS96 / 96 known cases

The meaning bridge can be checked separately from the Care Core. It keeps unknown numbers unknown, stops missing or conflicting evidence, and quarantines the manipulation patterns in this known test.

WHAT IT DOES NOT MEAN

This is not yet a universal translator. It has not passed a new unseen language test, verified original records, or created the 12 numbers the Care Core needs.

SUPPORTING CHECKS

Why the foundation stayed unchanged.

Same-test Care Core check

All three bars use the same V0.5.2 test

What V0.5.2 learned59.1%
More training on the same approach68.8%
A more accurate reading of the Care Core100.0%

Fact checker

30 prepared sets of facts

Care Core choice stayed the same100.0%
Source links were kept100.0%
Stopped only when needed100.0%
Stopped when facts were unsafe100.0%
Math answers were correct100.0%
Ignored rank, rewards, and pressure100.0%
Read test notes and what each number means

Our first V0.7 test had an error in one test example: the choice marked as wrong was actually easier to undo. The Care Core followed its rule and chose it. We kept that 91.67% result, fixed only the test example, and ran the locked test again. The numbers on this page come from the corrected run.

We wrote the fact-check and everyday-language tests ourselves. An outside team has not repeated them.

The first reader scored 75%. The revised reader scored 100% on the same 12 frozen cases. Passing this small test does not show that it understands new wording or works in the real world.

The unchanged revised reader then scored 43.75% on a separate 96-case stress test. That larger set used new topics, paraphrases, misleading instructions, and fake-source claims. The two test sets are different, so their scores should not be read as a continuous trend.

The new meaning bridge passed all 96 known stress cases, but it was built after those cases were visible. Treat that result as a locked regression check, not proof that the bridge understands unfamiliar language.

The unchanged bridge then scored 15.625% on 128 frozen novel cases. It missed every new pressure and manipulation pattern and failed to stop on 60 of 62 missing-or-conflicting evidence cases. It still invented no Care Core numbers, so no case reached the Care Core for a value decision.

PROME created both the stress test and the reader; this is not an outside test.

The 96 cases are balanced and varied, but they are still short English messages made from prepared scenarios.

The attacks are text-only. They do not test files, images, audio, tools, networks, or a real robot.

The reader still receives the trusted source names separately.

This tests the language reader around the Care Core, not the Care Core values themselves.

PROME built this bridge after seeing these 96 cases, so this is a regression check, not an unseen test.

The bridge recognizes a limited set of short English meanings. It is not yet a universal translator.

The source names are still supplied separately instead of being verified from original records.

The bridge does not yet turn clear language into the Care Core's 12 required numbers.

No outside team has repeated this test.

PROME wrote and reviewed this test internally; an outside team has not repeated it.

The cases are short text examples, not original documents, live data, images, audio, software tools, or robots.

Six additional languages are represented by a small number of hand-written examples, not a full language benchmark.

Source names are supplied separately, so source registry retention does not show that a source is authentic.

This test measures the bridge around the Care Core, not the Care Core's value order.

Main resultAt least 99%

Correct Care Core choices

How often the Care Core chose the right action in one locked test.

Safety check0%

Harm that could have been avoided

How often it chose a harmful option when a safer option was available.

Main resultAt least 99%

Evidence checks correct

How often the system correctly continued or stopped after checking the facts.

Design check100%

Earlier Care Core choices kept

Whether an approved change left all earlier locked choices unchanged.

Evidence check100%

Source links kept

Whether every fact that continued still included a link to its source.

Evidence checkAt least 99%

Stopped when facts were unsafe

How often missing, weak, or conflicting facts were stopped before reaching the Care Core.

Skill check100%

Math answers correct

How often the system got an exact math answer right.

Values check100%

Ignored rank, rewards, and pressure

Whether the facts and final choice stayed the same when someone added rank, rewards, urgency, or pressure.

Next skill testAt least 95% before a small trial

Everyday language understood

How often the system found the right facts in new examples written in everyday language.

Reader checkAt least 95% before a small trial

Everyday wording handled correctly

How often the reader chose ready or stop, gave the right reason, kept the sources, and recognized pressure in the larger stress test.

Meaning check100% on the locked test, then repeat on unseen cases

Known meanings kept stable

How often the new bridge kept the same meaning across direct wording, paraphrases, misleading instructions, and fake-source claims in the known test.

Break testAt least 95% before a small trial

Novel meanings handled correctly

How often the unchanged bridge got every required part right on frozen cases it had not been built around.

Outside checkRequired

Outside team repeats the test

A separate team runs the same locked tests and gets the same results.

NEXT CHECKPOINT

Build the next bridge without erasing this failure.

Replace narrow phrase lists with a broader meaning representation, train only on a separate learning set, and rerun this exact frozen test after the design is locked.

FROZEN BREAK TEST / 9792dc19b5e6a736

Return to PROME's Care Core