Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

4 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Redbelly DAO Proposal Pre-Screening Tool

Live: https://redbelly.smartcodedbot.com/screen/

Checks a draft proposal against the 11-criterion review checklist the Redbelly Community DAO ratified on 2025-10-03, and returns a per-criterion verdict with the authority behind it. Built for TASK-25 on the DAO Task Board.


Why this exists

The DAO already has a real, agreed standard for reviewing proposals. On 2025-10-03 it ratified "Proposal Review Checklist for Redbelly Community DAO" with 4,347,753 votes For and zero Against. Eleven criteria, each with pass conditions and a flag condition.

The problem is not that the standard is missing. It is that the standard lives in the body of a Snapshot proposal, and a person writing a draft has no way to run it against their work. They discover what the checklist says after the DAO has voted.

That is the gap this fills. Paste your draft in, find out what gets flagged and why, fix it, then submit.


What it does

Two layers, in this order

1. Deterministic structural checks run first. Whether a budget separates USDT from RBNT, whether milestones carry payment amounts, whether the nominated reviewer is also the proposer: these are facts about a submission, not opinions about it. They run as ordinary code. A council member can defend a flag from this layer in a meeting, because the reasoning is a comparison, not a model's judgment.

2. A language model judges only what structure cannot decide. Whether a KPI is meaningful. Whether a stated risk is really a risk. Whether a goal is vague. The model is given the ratified text of the criterion and told to judge against that text rather than against its own idea of a good proposal. It is instructed to return "unclear" rather than guess.

The model can never overturn a deterministic flag. It only fills gaps. This matters because the flag that catches the June 2026 FINPR proposal, a proposer naming themselves as their own paid "Independent Reviewer", must be provable rather than attributable to a model.

It fails loudly

If the model is unreachable, the structural checks still run and the report says which criteria went unassessed and why. It never returns fewer findings and presents that as a clean pass. An empty result is not success.

Ask by typing

A grounded question-answering panel reads Constitution v1.2, the ratified checklist, and all 30 proposals in the Snapshot archive. It answers from those documents and cites what it used, or says the answer is not in them. It is explicitly instructed to say when a rule has no constitutional basis, because three of the eleven criteria are in that position and a model asked about them will otherwise invent an anchor.


What the mapping turned up

Tracing each criterion back to its source surfaced things that are true but are not written down in either document. These are surfaced in the tool rather than hidden behind a clean checkmark.

Three criteria have no constitutional anchor at all

Criterion Finding
7. Risk and Mitigation The word "risk" does not appear anywhere in Constitution v1.2. There is no risk assessment, mitigation or disclosure requirement in the document.
8. Co-Funding and Leverage No co-funding, matched-funding or external-contribution clause exists. The criterion itself calls this "not a dealbreaker", so the tool treats it as advisory rather than blocking.
10. Compliance and Ethical Standards v1.2 contains no code of conduct, behavioural standards or enforcement section. The words "conduct", "behaviour" and "behavior" do not appear.

Criterion 10 is the clearest case of a rule living entirely outside the constitution: the standard it refers to was created by the "Community Code of Conduct" proposal, which passed For 93% on 2025-10-03. The tool cites that proposal as the authority, because there is no constitutional section to cite.

The task brief's own hint is wrong

The TASK-25 description suggests mapping budget limits to section 6.2. Section 6.2 Budgeting contains no percentage cap. It sets an annual cap of 100,000 USDT and a 10M RBNT incentive pool, and its only percentage runs the other way: a 10% minimum floor reserved for Open Innovation, which is a reservation, not a ceiling.

Criterion 1's requirement that a request "does not exceed 25% of the working group's annual allocation" has no constitutional basis anywhere. It is enforced by the ratified checklist alone. The tool says so on every result.

The live constitution has drifted from the ratified one, unversioned

There are two documents both labelled v1.2: the ratified PDF in the DAO Resources section, and a live Notion page. They differ, and the page has never been version-bumped.

Where Ratified PDF Live Notion page
Section 4.1 step 1 "Submit proposal via Google form" "Submit proposal via Notion form"
Section 4.1 step 5 "Vote via Snapshot" adds "or Discord (for holders with <5k RBNT)"
Section 6.2 one section, titled Budgeting a second 6.2, titled Transparency

The live page introduces a second voting venue and a wealth-based routing rule that was never ratified. Section 9 requires "Version control and public publishing required for each update", so this is a constitutional change made by document edit rather than by vote.

This tool parses the ratified PDF. Citing the live page would cite text nobody voted on. Every citation still carries the section title as well as the number, because a bare "6.2" is ambiguous across the two documents even though it is unambiguous within the ratified one.

Correction. An earlier version of this tool parsed the Notion page and reported, as a headline finding, that "Constitution v1.2 contains two sections numbered 6.2". That is true of the live page and false of the ratified text. The source is now the PDF, the divergence is reported as its own finding, and the test suite asserts the ratified text has exactly one 6.2 so the mistake cannot recur.

Five later votes override the constitutional text

The constitution was ratified 2025-09-17. Seventeen proposals have passed since, and five of them change how a criterion should be read:

Proposal Date Effect
Budget Allocation per Proposal 2025-09-17 For 81%. A single proposal may exceed one quarter's budget allocation.
Community Code of Conduct 2025-10-03 For 93%. Creates the standard Criterion 10 refers to.
Community Consultation Policy 2025-12-05 For 100%. The operative consultation rule, stronger than the constitution's single 3-day window.
Amendment to Leadership Structure 2026-07-11 For 97%. Changes the leadership structure oversight roles sit within.
Temporary Suspension of new treasury-funded activities 2026-06-29 For 100%. New treasury funding is paused until a future vote.

The suspension is applied as an overlay: any proposal requesting treasury funds gets told it sits under the suspension, regardless of whether its budget is otherwise compliant. That is a fact about the rules currently in force, not a defect in the proposal.

Three named working groups, five pods

The ratified criterion names three working groups (Community, Marketing, Developers/Builders). Constitution section 3 establishes five pods (Marketing, Builder/Develop, Researcher, Community, Partnerships). A proposal aligned to the Researcher or Partnerships pod satisfies the constitution but not the literal text of Criterion 3. Neither document resolves this.


The worked examples

The best test of a screening tool is a proposal whose answer is already known. The task names the right pair, and it is a good one: same author, same vendor, three months apart, opposite outcomes.

Proposal Date Outcome
Approved Marketing press only : FINPR Agency 2026-03-01 For 85.4%, zero Against
Rejected Continuation of Long-Term Marketing & PR Campaign with FINPR 2026-06-29 Against 71.4%, zero For

Run both from the Worked examples tab. The report is produced from the archived text with no knowledge of the outcome.

What the tool says

Approved proposal Rejected proposal
Flags (structural layer) 1 4
High severity 0 2
Full pipeline incl. model 1 7

The two high-severity flags are exactly the defect:

  • Criterion 5 Oversight. The nominated reviewer is the proposer. Constitution section 5: "Recusal is required for any vote on self-submitted proposals."
  • Criterion 9 Contribution Equity. One person holds two paid positions. Constitution section 5: "Limit: One paid role per individual per quarter."

Why this pair is a hard test, not a soft one

The rejected proposal is well built on almost every structural axis. It separates USDT from RBNT, ties every payment to a milestone, lists four risks each with a mitigation, and states KPIs. A shallow checker passes it. The single thing wrong is that the proposer pays themselves 2,400 USD equivalent as the "Independent Reviewer" of their own proposal, while the text asserts they hold no other paid DAO role.

So the bar is not "flag the bad one". It is: flag the conflict, and do not drown it in noise about the parts the proposal got right. The test suite asserts both directions, that criteria 1, 6 and 7 do not flag on the rejected proposal, because a checker that flags everything is as useless as one that flags nothing.

A finding that does not flatter the tool

Criterion 2 flags the approved proposal too, and the flag is correct. Its body requests "70% upfront in usdt" and then, four lines later, states "No upfront payments." That is a contradiction inside a proposal the DAO passed at 85.4%.

This was originally written as an expected pass, and the expectation was wrong, not the engine. The engine was not weakened to satisfy it. Passing a vote does not make a proposal checklist-clean, and a pre-screen that only ever agreed with the outcome would tell you nothing you did not already know.

The examples cannot be invented

engine/build-examples.py derives both from the archive, and every field value carries the exact quote from the proposal body that supports it. The script verifies each quote appears in the archived text and refuses to write the file if any does not. Fields recorded as absent are absent in the original. A reviewer can check any field by searching the body for its quote.


Running it day to day, for a council member

You do not need to install anything. Open https://redbelly.smartcodedbot.com/screen/.

Screening a proposal that has been submitted to you. Open the Screen a proposal tab and fill in what the proposal says. Fields you leave blank are treated as absent, which is exactly what the checklist does, so an incomplete proposal produces an accurate report rather than a forgiving one. Press Run the check.

Reading the report. Flags sort to the top, most serious first. Each one gives you the reason, the ratified flag condition it came from, and the authority behind it. Two labels matter:

  • structural check means the finding came from the submission's data. It is reproducible and you can defend it in a meeting.
  • model judgment means a language model assessed it. Weigh it accordingly, and read the quote it gives before acting on it.

Also watch for no constitutional anchor and partial anchor. These tell you a criterion is enforced by the checklist alone or rests on text a later vote has overtaken. That is not a reason to ignore the flag, but it is a reason to say so out loud if a proposer challenges it.

Answering a proposer who disagrees. Use the Ask tab. It answers from the constitution, the checklist and the archive, and cites what it used. If the rule they are asking about does not exist, it will say so rather than inventing one.

When a report says criteria went unassessed. The model layer was unreachable. The structural findings are complete and unaffected. The unassessed criteria need a human read. Nothing has been silently passed.

Keeping the mapping current. See the maintenance section below. In practice: after any governance vote that changes process, re-run the build and check the entries flagged softGround.


Nothing is hardcoded

Every criterion, constitutional section and proposal is generated from live sources by one script:

python3 engine/build-sources.py -o engine/
Output Source
checklist.json Snapshot GraphQL, the ratified proposal body, split into 11 criteria
constitution.json The ratified PDF, 17 sections with verbatim text, fetched from the DAO Resources API
archive.json Snapshot GraphQL, 30 proposals with outcomes
supersessions.json Derived: passed proposals dated after ratification
examples.json Derived: the two FINPR proposals, every field quote-verified against the archived body

Every artifact is stamped with the date it was built, and the report shows it. Serious citators all carry an "as of" date on a validity signal, because a signal with no date cannot be told apart from a stale one. The DAO amends its rules by vote, so a reader needs to know how current the reading is.

The criteria are parsed out of the ratified text, not retyped, so the pass and flag conditions are guaranteed to match what was actually voted on. The test suite asserts that the criterion titles match the ratified body exactly.

Exactly one file in the pipeline carries human judgment: engine/mapping.json, the map from each criterion to the constitutional section it enforces. It is kept separate from the generated data so a reviewer can disagree with the mapping without touching anything else. Every entry that rests on weak ground carries a softGround flag so the next maintainer knows where to look first.

Updating the mapping when the constitution is amended

  1. Re-run npm run build. This refreshes the criteria, the constitution, the archive, the supersessions and the worked examples from source.
  2. Run npm test. 57 assertions. If a criterion title, a section number or a quoted phrase changed, a test fails and names it. In particular the suite asserts the criterion titles still match the ratified text exactly, so a drifted parse cannot pass silently.
  3. Open engine/mapping.json and revisit every entry with softGround: true. That flag exists to tell the next maintainer where to look first. Check whether the amendment gave a previously unanchored criterion a real home, or moved a section number.
  4. If a governance vote changed process without amending the document, add it to that criterion's supersededBy with controlling: true and a plain-English effect. This is how the tool represents a rule that lives outside the constitution.
  5. Update mappingReviewed to today. The report shows this date.

The site renders from the JSON, so it cannot drift from the rules it checks.


Running it

npm test      # 65 assertions, deterministic layer plus worked examples
npm run build # regenerate all data from live sources
npm start     # serve on :3126

For the model layer, set FREEMODEL_API_KEY and VIRTUALS_API_KEY in .env. The tool runs without them and reports the affected criteria as unassessed.

Keys never reach the browser. All model calls go through the server. The client only ever sees the finished report.

Deploy with ./deploy.sh --restart, which checksum-verifies every file before moving it into place. That is not ceremony: a shell redirect truncates its target the moment it opens, so an interrupted transfer leaves an empty file behind, and an empty JS module imports cleanly while exporting nothing. It happened once during this build. The server now also refuses to start if any module failed to load, turning a silent runtime failure into a loud startup one.


Limits, stated plainly

  • It is a pre-screen, not a decision. It reports what the ratified checklist says about a draft. It has no vote and no standing. A proposal it passes can still be rejected. A proposal it flags may be entirely reasonable once explained.
  • It reads one proposal at a time. It cannot see patterns across proposals. In the worked example, the same wallet filed the treasury-suspension proposal nine minutes before the FINPR continuation. No checklist criterion covers that, and the tool cannot surface it. That is a limit of the approach.
  • Identity matching is textual. It compares names and handles after stripping case and punctuation. Someone determined to disguise a conflict by using a different name will not be caught.
  • The model layer is judgment, not fact. Results from it are labelled as such in the report so a reviewer can weigh them accordingly.
  • The mapping is one person's reading. It is published in full, with every weak anchor called out, precisely so it can be argued with.

Sources

MIT.

About

Automated pre-screening for Redbelly DAO proposals, checked against the 11-criterion review checklist ratified 2025-10-03. Every result cites the criterion and the constitutional section behind it.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages