A future comparison
A useful study could compare traditional prompts with prompts that use a documented Gaiish structure across representative tasks and models. The comparison would need the same task brief, source material, model conditions and evaluation process, with enough varied examples to avoid treating one prompt or one model as representative of all use. The design, sampling, exclusions and analysis would need to be documented before results were interpreted.
Measures to define
| Measure | What a study could examine |
|---|---|
| Accuracy | Whether factual or calculated claims match the supplied evidence or reference answer. |
| Completeness | Whether the response covers the required information without material omissions. |
| Instruction adherence | Whether the response follows the task and stated requirements. |
| Formatting compliance | Whether the response matches the requested structure, schema or limits. |
| Hallucination frequency | How often the response introduces unsupported facts, sources or details. |
| Human preference | Which response reviewers prefer under a defined rubric and task context. |
| Iterations required | How many prompt–response revisions are needed to reach a predeclared criterion. |
What would need controlling
Models differ in training, context limits, system instructions, tools and response style. A study would need to record the model and version, prompt text, supplied material, temperature or equivalent settings where available, evaluator rubric and human-review procedure. It would also need to separate the effect of explicit structure from extra words or extra context.
What this page does not say
There is no published result here, no claimed percentage improvement and no claim that one framework is best for every task. Structure can improve instruction adherence and reduce ambiguity as a practical hypothesis; whether it does so under a particular condition is an empirical question.
Read the methodSee the language specificationAbout the project