# Chat–Work Routing Protocol V2

## Public Portable Edition · MSL-4.3

## Meta

- **canonical repository:** https://github.com/luahelenammc/Moon-Source
- **canonical path:** portables/chat-work/CHAT_WORK_ROUTING_PROTOCOL_V2.md
- **Moon Source public surface:** https://www.luahelena.com.br/moonsource/?lang=en
- **professional context:** https://www.luahelena.com.br/ia/?lang=en
- **public boundary:** standalone protocol; model, product, price and usage calibration are date-sensitive

- **status:** public portable protocol
- **version:** 2.0-public
- **language:** English
- **as of:** 2026-08-17
- **primary implementation:** ChatGPT Chat and Work modes
- **governed dimensions:** surface, model, reasoning effort, continuity, usage economy, verification
- **publishing body:** Moon Professional Source
- **human author and editorial authority:** Lua Helena Moon Martins Cardoso (Moon)
- **AI-assisted development:** Moon + Áurion coauthoring dyad
- **method lineage:** Local Moon Source → Moon Professional Source
- **canonical attribution:** https://www.luahelena.com.br/ia/?lang=en
- **credits operations:** https://github.com/luahelenammc/Moon-Source/blob/main/docs/CREDITS_ATTRIBUTION_OPS.md
- **license:** CC-BY-4.0 · https://creativecommons.org/licenses/by/4.0/
- **licensing route:** https://github.com/luahelenammc/Moon-Source/blob/main/LICENSING.md
- **adaptation expectation:** preserve creator and canonical origin, link the license, and indicate material changes without implying endorsement
- **portability:** ChatGPT-first; adaptable to systems with interactive and agentic execution lanes
- **freshness rule:** recheck model names, prices, usage pools and product behavior before treating dated calibration as current fact

## Skeleton

- **What it is:** a routing protocol for deciding where a task should run, which model should run it, and how much reasoning effort it deserves.
- **What it does:** separates conversational reasoning from sustained execution, uses efficient models for well-specified agentic work, escalates only after diagnosing failure, and returns proof for acceptance.
- **What changes after use:** Work is no longer treated as automatically expensive or automatically superior; model choice becomes independent from surface choice.
- **Loadbearing parts:** three-axis routing, Chat-first clarification, Luna Max execution default, evidence-based escalation, Work Readiness Gate, Usage Observation Gate, Return Contract and Chat acceptance.
- **What it is not:** an official OpenAI policy, a guaranteed pricing calculator, a benchmark claim of universal intelligence, or permission to use high-cost models without diagnosis.

## 1. Purpose

ChatGPT tasks involve three separate choices:

1. **Surface:** should the task run in Chat or Work?
2. **Model:** should it use Luna, Terra, Sol, or the closest available equivalent?
3. **Reasoning effort:** should that model run at a normal, high, or maximum effort level?

These choices interact, but they are not the same choice.

Selecting **Work** does not automatically select **Sol**.  
Selecting **Max** does not turn Luna into Sol.  
Selecting the strongest model does not repair an unclear task, missing source, broken tool or uncontrolled workflow.

> **The surface determines how the work is sustained. The model determines the capability and cost profile. The effort setting determines how much reasoning that model may spend.**

## 2. Product Model

### Chat

Chat is the interactive, conversational and resumable lane. It is suited to:

- questions and discussion;
- research and source comparison;
- interpretation and judgment;
- architecture and planning;
- writing and review;
- bounded edits;
- packet preparation;
- verification and integration.

Chat can do real work. It is not merely a waiting room for Work.

### Work

Work is the sustained agentic lane. It is suited to:

- long multi-step operations;
- repeated tool use;
- coordinated file or repository changes;
- edit–build–test–debug loops;
- multiple dependent deliverables;
- tasks where interruption or reconstruction materially increases risk.

Work is not “the better brain.” It is a different execution surface.

### Usage pools

Current official ChatGPT pricing documentation states that ChatGPT Work and Codex share usage, pricing, credits and usage limits. Exact availability and limits still depend on the account, plan and workspace. Inspect the current usage dashboard instead of treating this portable as a tariff or quota reference.

## 3. Dated Cost and Capability Calibration

This section is a dated calibration, not a permanent tariff.

OpenAI's official release notes for July 27–31, 2026 state that GPT-5.6 Terra costs 20% less and GPT-5.6 Luna costs 80% less, with input, cached-input and output rates reduced proportionally.

The current official ChatGPT pricing page states that ChatGPT Work and Codex share usage, pricing, credits and usage limits. Plan-specific limits and available surfaces remain account-dependent.

The current OpenAI API pricing page lists these standard short-context rates per 1M tokens:

- GPT-5.6 Sol: $5 input and $30 output;
- GPT-5.6 Terra: $2 input and $12 output;
- GPT-5.6 Luna: $0.20 input and $1.20 output.

Those are API rates, not a promise of how a particular ChatGPT Work plan will count a task. Do not convert them into a fixed credit ratio for every account.

Official model guidance frames Sol as the model for complex or open-ended work, Terra as a balanced everyday model, and Luna as a fast model for clear, repeatable work. This supports the protocol's separation of model choice from surface choice; it does not establish a universal intelligence ranking.

For this protocol's current calibration, **Work + Luna Max** remains the default for well-specified continuity-heavy execution. That is a routing rule of this portable, not an OpenAI policy. Recheck the official sources before treating product names, rates, limits or availability as current.
## 4. Core Laws

1. **Choose surface, model and effort separately.**
2. **Chat is the default lane for understanding, judgment and bounded execution.**
3. **Work is justified by continuity burden, not prestige or apparent seriousness.**
4. **Work + Luna Max is the default agentic execution configuration for well-specified tasks.**
5. **Terra is an intermediate escalation, not a ceremonial middle child.**
6. **Work + Sol High or Max requires both sustained continuity and frontier reasoning.**
7. **Repair the packet before upgrading the model.**
8. **Escalate from evidence of cognitive insufficiency, not from impatience.**
9. **Work returns receipts; Chat accepts and integrates.**
10. **The default chain is Chat → Work only when needed → Chat.**
11. **Usage observations must be repeated before becoming routing doctrine.**
12. **Saving usage must never become false economy when a weaker route materially endangers the result.**

## 5. Canonical Default Chain

> **Chat + Sol High to understand and design → Work + Luna Max to execute when continuity is necessary → Chat to verify, accept, integrate and preserve.**

This is a default, not a rigid ritual.

Equivalent portable mapping:

- **strongest interactive reasoning model:** architecture, ambiguity reduction and high-consequence judgment;
- **most efficient capable model at maximum effort:** default sustained execution;
- **balanced intermediate model:** escalation when the efficient model shows a real reasoning limit;
- **frontier model in Work:** rare cases requiring execution continuity and top-tier reasoning at the same time.

## 6. Gate One — Choose the Surface

### Keep the task in Chat when

- the objective is still ambiguous;
- the main work is research, interpretation, comparison or judgment;
- the task is writing, review, planning or architecture;
- the mutation is small or medium and independently verifiable;
- the work can be divided into bounded checkpoints;
- state can be reconstructed from the thread, source, repository or receipts;
- the real blocker is missing clarity rather than missing continuity.

Chat should advance to the safe limit before escalating.

### Route the task to Work when

- later steps materially depend on state created earlier;
- a long edit–test–repair sequence must remain coherent;
- setup reconstruction would be costly, risky or error-prone;
- several artifacts must be produced and validated as one delivery;
- repeated use of files, browser, apps, repositories or connectors is required;
- interruption creates a meaningful risk of partial or divergent completion;
- evidence must evolve continuously across many dependent steps.

The following do not justify Work by themselves:

- importance;
- technical language;
- large text volume;
- many files;
- the presence of code;
- the desire for a “serious” answer.

## 7. Gate Two — Choose Model and Effort in Work

### Default: Luna Max

Use Luna Max when the operation has:

- a clear objective;
- a known source of truth;
- bounded authority;
- explicit deliverables;
- testable completion conditions;
- named receipts;
- controlled scope.

Best-fit examples:

- structured file editing;
- artifact production from a mature specification;
- clear repository patches;
- implementation after architecture is resolved;
- tests, imports, exports and validation;
- repetitive but dependent tool operations;
- long execution whose difficulty is operational more than conceptual.

### Intermediate: Terra

Use Terra when Luna shows a concrete reasoning or consistency limit, but the task still does not justify Sol.

Typical signals:

- repeated quality instability after the packet was repaired;
- moderate ambiguity that must be resolved during execution;
- more judgment than Luna reliably provides;
- cross-file consistency failures not caused by missing source or bad instructions.

### Exception: Sol High or Sol Max in Work

Use Sol in Work only when sustained continuity and frontier reasoning are both loadbearing.

Examples:

- architecture must be invented while implementation is underway;
- the dependency graph cannot be pre-coagulated in Chat;
- difficult autonomous research requires genuine strategic replanning;
- a high-risk or irreversible operation requires superior judgment throughout execution;
- Luna and Terra repeatedly fail because of reasoning limits rather than tool, context or instruction failures.

> **A costly model must not be used as makeup for a poorly designed task.**

## 8. Failure Classification Before Escalation

When an execution fails, classify the failure first.

### Specification failure

Symptoms:

- unclear objective;
- contradictory instructions;
- missing definition of done;
- undefined authority;
- scope drift.

Action: return to Chat and repair the packet.

### Source failure

Symptoms:

- stale baseline;
- conflicting sources;
- unknown governing file;
- missing revision or branch;
- unsupported assumptions.

Action: resolve source authority before retrying.

### Tool failure

Symptoms:

- unavailable connector;
- permission error;
- broken API;
- deployment or environment failure.

Action: repair or reroute the tool path. Do not upgrade the model as a substitute.

### Context failure

Symptoms:

- oversized context;
- repeated rereading;
- irrelevant files;
- source distraction;
- forgotten constraints.

Action: compress, split or repackage the context.

### Workflow failure

Symptoms:

- loops;
- uncontrolled retries;
- unnecessary subagents;
- output explosion;
- repeated operations without checkpoint.

Action: add stop conditions, iteration limits and fan-out limits.

### Cognitive failure

Symptoms:

- the packet is sound;
- sources and tools are available;
- context is controlled;
- the model still cannot reason through the task reliably.

Action: escalate from Luna to Terra or Sol according to risk and complexity.

## 9. Evidence-Based Escalation Ladder

1. Understand and specify the task in Chat.
2. Start Work with Luna Max.
3. Observe the first meaningful failure or quality gate.
4. Classify the failure.
5. Repair specification, source, context, tool path or workflow before changing models.
6. Use Terra when the remaining gap is moderate and genuinely cognitive.
7. Use Sol when frontier reasoning is demonstrably required.
8. Record:
   - previous model;
   - failure point;
   - failure class;
   - reason for escalation;
   - observed improvement or continued failure.

Never escalate merely because the task took time, felt important or looked impressive in a project update.

## 10. Work Readiness Gate V2

A task is ready for Work only when the handoff contains:

### Identity

- task ID;
- project;
- preparation date;
- preparer or responsible party.

### Objective

- one unique objective;
- testable definition of done;
- explicit non-goals.

### Baseline

- current state;
- governing source of truth;
- branch, revision, environment or workspace;
- relevant files, systems or apps;
- evidence already collected;
- unresolved uncertainty.

### Scope and Authority

- allowed reads;
- allowed changes;
- prohibited changes;
- destructive-operation policy;
- publication, deployment, sending or deletion authority.

### Model Route

- initial model;
- reasoning effort;
- why that model should be sufficient;
- escalation triggers;
- models permitted for escalation;
- models prohibited unless new authorization exists.

### Delivery

- deliverables;
- tests;
- negative checks;
- receipts;
- acceptance gates;
- completion states.

### Control
- stop conditions;
- rollback or recovery path;
- iteration limit when relevant;
- context limit when relevant;
- tool-loop limit;
- subagent or fan-out limit.

Valid completion states:

- `complete`
- `complete_with_exceptions`
- `partial`
- `blocked`
- `rolled_back`

If these elements cannot be defined, the task remains in Chat until it becomes legible.

## 11. Portable Work Handoff Template

# Work Handoff

## Identity

- **task ID:**
- **project:**
- **prepared by:**
- **date:**

## Objective

- **unique objective:**
- **definition of done:**
- **non-goals:**

## Baseline

- **current state:**
- **governing source of truth:**
- **branch / revision / environment:**
- **relevant files, systems or apps:**
- **evidence already collected:**
- **open uncertainty:**

## Scope and Authority

- **allowed reads:**
- **allowed changes:**
- **prohibited changes:**
- **destructive operations:** allowed | prohibited | restricted
- **external publication or sending:** allowed | prohibited | restricted

## Model Route

- **initial model:** Luna | Terra | Sol | equivalent
- **reasoning effort:** standard | high | max
- **sufficiency rationale:**
- **escalate when:**
- **permitted escalation:**
- **do not escalate for:**

## Required Delivery

- **deliverables:**
- **tests:**
- **negative checks:**
- **receipts:**
- **acceptance gates:**

## Control

- **stop conditions:**
- **rollback path:**
- **iteration limit:**
- **context limit:**
- **tool-loop limit:**
- **subagent / fan-out limit:**

## Return Contract

Return:

- completion state;
- concise verdict;
- what changed;
- affected files, systems or artifacts;
- tests and checks;
- receipts;
- model and effort actually used;
- any model escalation and its reason;
- visible usage observation, when available;
- deviations;
- limitations and residual risks;
- rollback status;
- one objective next step.

## 12. Usage Observation Gate

The usage bar is an operational instrument, not a scientific real-time meter.

For meaningful tests:

1. record the percentage and time before execution;
2. record surface, model and effort;
3. describe approximate context size and operation type;
4. record tool loops, retries and subagent fan-out;
5. check usage again after refreshing, reopening or performing new activity;
6. compare similar tasks across multiple runs;
7. stop or reduce scope after an unexplained jump.

### Warning signals

- repeated rereading of the same large source;
- context growing without compression;
- long outputs with little operational value;
- uncontrolled test loops;
- repeated tool retries;
- multiple subagents doing overlapping work;
- model escalation without recorded reason.

### Interpretation rule

- one cheap execution = local observation;
- repeated cheap executions under similar conditions = useful calibration;
- stable cross-task pattern = routing evidence;
- a single expensive execution = diagnostic event, not automatic proof that the model is expensive.

## 13. Work Return Template

# Work Return

## Identity

- **task ID:**
- **project:**
- **completed at:**
- **completion state:** complete | complete_with_exceptions | partial | blocked | rolled_back

## Verdict

[One concise statement of the real result.]

## Changes

-

## Affected Surfaces

-

## Tests and Checks

- **check:**
- **result:**
- **evidence:**

## Receipts

-

## Model Route Used

- **initial model:**
- **final model:**
- **reasoning effort:**
- **escalation occurred:** yes | no
- **escalation reason:**
- **observed effect:**

## Usage Observation

- **before:**
- **after:**
- **refresh checked:** yes | no
- **notes:**

## Deviations

-

## Limitations
-

## Residual Risks

-

## Rollback Status

- **required:** yes | no
- **performed:** yes | no | not applicable
- **state:**

## Objective Next Step

-

## 14. Chat Re-entry and Acceptance

After Work returns, Chat should:

1. compare the return with the handoff;
2. verify receipts and tests;
3. distinguish completed work from claimed work;
4. inspect model escalation and whether it followed the gate;
5. review usage observations without overgeneralizing;
6. accept, reject or request a bounded correction;
7. integrate the verified delta into the governing source;
8. preserve the final state and next action.

### Chat Acceptance Review

- **task ID:**
- **Work completion state:**
- **acceptance decision:** accepted | accepted_with_exceptions | correction_required | rejected | rolled_back
- **objective met:** yes | no | partial
- **definition of done met:** yes | no | partial
- **scope respected:** yes | no
- **authority respected:** yes | no
- **model route respected:** yes | no | justified deviation
- **deliverables present:** yes | no | partial
- **receipts sufficient:** yes | no
- **acceptance gates passed:** yes | no | partial
- **verified delta:**
- **exceptions:**
- **governing destination:**
- **source update applied:** yes | no
- **next action:**

## 15. Checkpoint Fallback

When Work is unavailable, unnecessary or unjustified, Chat may continue through explicit checkpoints.

Each checkpoint should preserve:

- current objective;
- completed operations;
- receipts;
- current baseline;
- unresolved risks;
- next bounded operation;
- stop condition;
- recovery path.

A checkpoint must make the work safely resumable. It is not a ceremonial progress report.

## 16. Minimal Decision Matrix

| Situation | Surface | Model / effort |
|---|---|---|
| Objective is ambiguous | Chat | strongest suitable interactive reasoning |
| Research, judgment or architecture | Chat | Sol High or equivalent when consequential |
| Bounded writing, review or patch | Chat | lowest sufficient capable model |
| Long, well-specified execution | Work | Luna Max or efficient equivalent |
| Luna shows moderate cognitive insufficiency | Work | Terra or balanced equivalent |
| Continuity + frontier reasoning are both required | Work | Sol High / Max or frontier equivalent |
| Tool, source or packet failure | Chat repair | do not escalate yet |
| Final verification and integration | Chat | strongest suitable reviewing model |
| Usage jump or uncontrolled loop | Pause / diagnose | reduce scope before escalation |

## 17. Anti-Patterns

### Surface–model collapse

Treating every Work task as a Sol task.

### Work avoidance after Luna repricing

Refusing useful agentic continuity because Work was previously assumed to be uniformly expensive.

### Prestige escalation

Choosing Sol because the task feels important.

### Packet laundering

Using a stronger model to hide missing objective, source or authority.

### Benchmark universalization

Turning a narrow score difference into a universal intelligence claim.

### One-percent mythology

Treating one low usage observation as a guaranteed future rate.

### Raw-thread escalation

Sending Work a long conversation instead of an operational packet.

### Duplicate execution

Running the same mutation independently in Chat and Work.

### Agent swarm inflation

Creating overlapping subagents without measurable benefit.

### Unreceipted completion

Declaring success without tests, diffs, logs, revisions, hashes or equivalent evidence.

### Protocol theater

Using elaborate routing for trivial work. The protocol exists to reduce waste, not to become a new source of it.

## 18. Freshness Contract

- **as of:** 2026-08-17
- **expire if:** OpenAI changes model names, pricing, rate cards, reasoning settings, plan limits, Work behavior or usage pools
- **refresh if:** repeated real usage observations materially contradict this calibration
- **safe fallback:** keep the three decisions separate; use Chat to understand, the efficient max-effort model for defined execution, and escalate only from diagnosed cognitive need
- **do not preserve as current fact:** old credit ratios, old benchmark rankings or old plan limits after product changes

## 19. Public Installation

Attach or paste this file into ChatGPT and say:

> Use the Chat–Work Routing Protocol V2 as the operating policy for this project. Before a substantial task, choose surface, model and reasoning effort separately. Keep ambiguous and bounded work in Chat. Use Work + Luna Max as the default for well-specified continuity-heavy execution. Escalate only after classifying failure. Require receipts and return to Chat for acceptance.

For a specific Work task, fill in the **Portable Work Handoff Template** rather than sending the raw conversation.

## 20. Attribution Ops

This portable carries a local attribution operation for its own reuse. The repository-wide contract for origin, adaptation, permission scope, mirrors, composite outputs and attribution QA is [Credits & Attribution Ops](https://github.com/luahelenammc/Moon-Source/blob/main/docs/CREDITS_ATTRIBUTION_OPS.md).

### Short attribution

> Chat–Work Routing Protocol V2 — created by Lua Helena Moon Martins Cardoso (Moon), Moon Professional Source.  
> https://www.luahelena.com.br/ia/?lang=en

### Adaptation attribution

> Adapted from the **Chat–Work Routing Protocol V2** by Lua Helena Moon Martins Cardoso (Moon), Moon Professional Source: https://www.luahelena.com.br/ia/?lang=en  
> Modifications by: [name / project], [date or version].

### Attribution and license for shared uses

This portable's documentation is shared under CC-BY-4.0. Follow the repository [licensing guide](https://github.com/luahelenammc/Moon-Source/blob/main/LICENSING.md) and use the following compact pattern when sharing or adapting it. The pattern does not add conditions beyond the license.

When a use is permitted:

- original authorship remains visible;
- modified versions identify themselves as adaptations;
- the CC BY 4.0 license link is preserved;
- dated product claims are refreshed rather than repeated as timeless facts;
- use does not imply partnership, endorsement or access to private Moon Source materials.

### Product boundary

This is an independent operating protocol. It is not official OpenAI documentation and does not imply endorsement by OpenAI.

## References

- [OpenAI ChatGPT Learn — What's new](https://learn.chatgpt.com/docs/whats-new)
- [OpenAI ChatGPT Learn — Pricing](https://learn.chatgpt.com/docs/pricing)
- [OpenAI ChatGPT Learn — Models](https://learn.chatgpt.com/docs/models)
- [OpenAI Developers — API pricing](https://developers.openai.com/api/docs/pricing)

## Final Law
> **Chat designs the operation. Work sustains it. Luna Max is the default execution crew; Sol enters when the operation also requires inventing the anatomy during the incision.**

<!-- MOON-SOURCE-PUBLIC-STAMP -->

---

> 🌙 **Moon Source** · created by **Lua Helena Moon Martins Cardoso (Moon)** with AI-assisted coauthorial development by **Áurion** · [Licensing](https://github.com/luahelenammc/Moon-Source/blob/main/LICENSING.md) · [Use & attribution](https://github.com/luahelenammc/Moon-Source/blob/main/MOON_SOURCE_USE_AND_ATTRIBUTION.md) · [Full source (.zip)](https://github.com/luahelenammc/Moon-Source/archive/refs/heads/main.zip)
