Correct Care Core choices
1,080 corrected choices after the Care Core was locked
GOAL / AT LEAST 99.0%CARE CORE RESULTS / METRICS-V0.6.0
What passed, what failed, and what PROME has not tested yet.
CURRENT RESULTS
1,080 corrected choices after the Care Core was locked
GOAL / AT LEAST 99.0%Corrected V0.7 locked test
GOAL / NO MORE THAN 0.0%30 prepared sets of facts; everyday wording tested separately
GOAL / AT LEAST 99.0%96 varied English stress tests; 42 fully correct
GOAL / AT LEAST 95.0%128 frozen novel cases; 20 fully correct
GOAL / AT LEAST 95.0%Our tests are evidence, not proof of real-world safety.
EXECUTIVE SUMMARY / FROZEN NOVEL BREAK TEST / MEANING-BREAK-RUN-V0.2.0
128 frozen cases · bridge unchanged · first result preserved · 20 fully correct
Different frozen tests; shown as a comparison, not a trend
Every bar starts at zero and uses all cases of that type
New wording, attacks, languages, and sentence traps
60 of 62 evidence stops were missed
0 of 31 pressure cases
0 of 49 attack cases
The bridge invented none of the 12 Care Core numbers. Even when it failed to stop bad evidence, every case still required numeric grounding and none reached the Care Core for a value decision.
The V0.1 bridge recognizes prepared phrase patterns, not meaning in a broad and reusable way. Its 95.8% stability score mostly means it repeated the same answer—even when that answer was wrong.
RESULTS BY VERSION
Different versions used different tests. Only V0.6 and V0.7 used the same locked test.
LANGUAGE STRESS TEST / EVERYDAY-READER-V0.2.0
96 varied English cases · 42 fully correct · first stress result preserved
Only the first two bars use the same test
Twenty-four cases in each bar
Twenty-four cases in each bar
35 of 48 required stops
Misleading instructions and fake sources
All 96 cases
The reader kept every source, but its fixed phrase lists did not understand many unfamiliar ways of saying that a fact was missing or that two sources disagreed. It incorrectly continued in 35 of 48 cases that should have stopped. Quoted attack words and copied source names also caused six false stops.
BUILDING TOWARD REAL USE
Right in all 8 focused V0.7 tests
Right in all 1,080 corrected tests
Right in all 30 home, business, and factory tests
42/96 fully correct; 35 of 48 required stops were missed
20/128 fully correct; the known 96-case contract did not generalize
Needed before we make claims about real-world use
Home, business, factory, and robot trials have not started
NEW MEANING BRIDGE / CARE-CORE-MEANING-V0.1.0
96 known cases · 24 base meanings · Care Core unchanged
Keep the trusted facts and source names.
Mark missing facts, disagreements, pressure, and manipulation.
Leave every unsupported Care Core number unknown.
A regression check built after these cases were visible
The meaning bridge can be checked separately from the Care Core. It keeps unknown numbers unknown, stops missing or conflicting evidence, and quarantines the manipulation patterns in this known test.
This is not yet a universal translator. It has not passed a new unseen language test, verified original records, or created the 12 numbers the Care Core needs.
SUPPORTING CHECKS
All three bars use the same V0.5.2 test
30 prepared sets of facts
Our first V0.7 test had an error in one test example: the choice marked as wrong was actually easier to undo. The Care Core followed its rule and chose it. We kept that 91.67% result, fixed only the test example, and ran the locked test again. The numbers on this page come from the corrected run.
We wrote the fact-check and everyday-language tests ourselves. An outside team has not repeated them.
The first reader scored 75%. The revised reader scored 100% on the same 12 frozen cases. Passing this small test does not show that it understands new wording or works in the real world.
The unchanged revised reader then scored 43.75% on a separate 96-case stress test. That larger set used new topics, paraphrases, misleading instructions, and fake-source claims. The two test sets are different, so their scores should not be read as a continuous trend.
The new meaning bridge passed all 96 known stress cases, but it was built after those cases were visible. Treat that result as a locked regression check, not proof that the bridge understands unfamiliar language.
The unchanged bridge then scored 15.625% on 128 frozen novel cases. It missed every new pressure and manipulation pattern and failed to stop on 60 of 62 missing-or-conflicting evidence cases. It still invented no Care Core numbers, so no case reached the Care Core for a value decision.
PROME created both the stress test and the reader; this is not an outside test.
The 96 cases are balanced and varied, but they are still short English messages made from prepared scenarios.
The attacks are text-only. They do not test files, images, audio, tools, networks, or a real robot.
The reader still receives the trusted source names separately.
This tests the language reader around the Care Core, not the Care Core values themselves.
PROME built this bridge after seeing these 96 cases, so this is a regression check, not an unseen test.
The bridge recognizes a limited set of short English meanings. It is not yet a universal translator.
The source names are still supplied separately instead of being verified from original records.
The bridge does not yet turn clear language into the Care Core's 12 required numbers.
No outside team has repeated this test.
PROME wrote and reviewed this test internally; an outside team has not repeated it.
The cases are short text examples, not original documents, live data, images, audio, software tools, or robots.
Six additional languages are represented by a small number of hand-written examples, not a full language benchmark.
Source names are supplied separately, so source registry retention does not show that a source is authentic.
This test measures the bridge around the Care Core, not the Care Core's value order.
How often the Care Core chose the right action in one locked test.
How often it chose a harmful option when a safer option was available.
How often the system correctly continued or stopped after checking the facts.
Whether an approved change left all earlier locked choices unchanged.
Whether every fact that continued still included a link to its source.
How often missing, weak, or conflicting facts were stopped before reaching the Care Core.
How often the system got an exact math answer right.
Whether the facts and final choice stayed the same when someone added rank, rewards, urgency, or pressure.
How often the system found the right facts in new examples written in everyday language.
How often the reader chose ready or stop, gave the right reason, kept the sources, and recognized pressure in the larger stress test.
How often the new bridge kept the same meaning across direct wording, paraphrases, misleading instructions, and fake-source claims in the known test.
How often the unchanged bridge got every required part right on frozen cases it had not been built around.
A separate team runs the same locked tests and gets the same results.
NEXT CHECKPOINT
Replace narrow phrase lists with a broader meaning representation, train only on a separate learning set, and rerun this exact frozen test after the design is locked.
FROZEN BREAK TEST / 9792dc19b5e6a736
Return to PROME's Care Core