# Testing the program and the decisions inside it

Design proposal · 7 September 2026 · for investigator, clinical and statistician review

## Recommendation

Use a staged program of research: observe families using the complete experience, optimize the expensive delivery components, then test the selected package against an appropriate comparison. Do not try to answer efficacy, kit contents, coaching dose, parent co-play, reminder timing, loss-specific tailoring and long-term maintenance in one underpowered trial.

The first optimization experiment I would prioritize is **2 × 2: physical kit provision × proactive parent coaching**. Both are consequential delivery/business decisions. All families receive the same game, reviewed parent discussion content, basic orientation, digital materials, and required clinical support. This trial would tell you the added value of the kit and proactive coaching within that program. It would not establish the efficacy of the game or the causal contribution of having any parent conversation.

## What the existing evidence does and does not support

The MGT open trial enrolled 65 youth; 42 completed Phase I and reported reductions across grief domains and other symptoms. Phase I included psychoeducation, skills and managing reminders. Its single-group design supports a rationale for this work, but does not isolate psychoeducation, estimate the effect of a home game, or establish maintained benefits for this adaptation. The authors also identify the need for further evaluation of maintenance. [Hill et al., MGT pilot open trial](https://mmhpi.org/wp-content/uploads/2021/12/Multidimensional-Grief-Therapy_Intervention-for-Bereaved-Children-and-Adolescents.pdf).

Family participation is a plausible mechanism, not merely a retention device. The Family Bereavement Program has randomized long-term evidence involving parenting and youth outcomes, including grief outcomes. It is a different intervention, focused on parental bereavement; it does not establish that discussion cards or coaching in Grief Galaxy will reproduce those effects across every loss relationship. [Family Bereavement Program six-year grief outcomes](https://pmc.ncbi.nlm.nih.gov/articles/PMC2874830/).

## Stage 0: make the experience and delivery model usable

Run an initial approximately 12–20-household usability study, purposively including younger children and teenagers, multiple caregivers, different loss relationships, low digital confidence, and both mailed and digital materials. This is an illustrative learning sample, not a powered clinical trial.

Observe a child playing, a parent preparing, and—with specific permission—a practice conversation or a role-play. Ask what the child wanted to share, whether the parent felt able to listen, which directions were misunderstood, and whether rewards felt appropriate. Test a late kit, a missed week, and a request for a call. Record staff time and contact attempts. Use short interviews to distinguish disinterest from practical barriers or discomfort.

Then run a small operational cohort through all six weeks and scheduled follow-up. Freeze the intervention and define feasibility criteria before deciding to proceed. Only three sessions are currently built; a three-week prototype pilot must be described as such and must not be treated as a test of the planned six-week program.

## Stage 1: optimize kit and coaching

### Common core

- Same six game modules and recreation opportunities, with the same content/build versions.
- Same brief orientation and parent listening guidance.
- Same basic weekly parent–child task, digital calendar/intentions tool and reminders.
- Same digital materials and printable files.
- Same safety assessment, clinical review, referral access, and on-request help appropriate to the service.
- Same scheduled research contacts, outcome measures, and assessment compensation.

### Factors

| Factor | Lower level | Higher level | What the contrast means |
|---|---|---|---|
| A: physical kit | Digital materials, including printable equivalents | Same digital access plus a standardized mailed/pickup kit: numbered envelopes, calendar and rewards | Incremental effect of proactively providing this physical bundle |
| B: proactive parent coach | Common orientation, usual technical/on-request support, and common clinical care | Same common support plus scheduled brief calls in weeks 1, 3 and 5 using a manualized support protocol | Incremental effect of proactive coaching under these support conditions |

This produces four combinations. For each main effect, compare everyone assigned the higher level with everyone assigned the lower level, averaging over the other factor. Do not power the trial as four separate pairwise comparisons. A kit-by-coaching interaction is clinically plausible: materials may work better when someone helps parents use them. Pre-specify whether that interaction is confirmatory or exploratory and size the study accordingly.

This follows the distinction between component optimization and later package evaluation in MOST. Factorial and sequential designs answer different questions; neither is a shortcut around adequate sample size. [NIH research-methods overview](https://researchmethodsresources.nih.gov/methods/additional).

The kit experiment tests a bundle, not the separate contribution of stickers, envelopes, calendar placement or mailing frequency. Keep those fixed during this experiment. Record control-arm self-printing, receipt, kit refusal and actual use. If families can decline a shipment, define the factor as offer/provision policy and analyze the randomized assignment; never silently reclassify decliners into the digital group.

### What to do about the parent discussion itself

My default is to retain a brief discussion task in the common core because it is central to the intended family program. That is a design choice, not a proven active ingredient.

If the grant's central question is instead whether **structured caregiver practice** matters, add or substitute a factor: basic conversation invitation versus a structured card with listening rehearsal and a small shared practice. A 2 × 2 × 2 design yields eight combinations. Choose this only with a clear primary question and sufficient recruitment; it does not independently answer whether all parent involvement can be removed. If recruitment is limited, prioritize two factors rather than calling a small eight-condition study definitive.

Keep parent co-play versus independent play a documented family preference initially, especially across ages 8–16. A later experiment could randomize an offered co-play format where both options are acceptable. Observational differences between families choosing co-play and solo play would not establish which is better.

### Population, unit and setting

Enroll bereaved children aged 8–16 with the elevated grief-related difficulties specified by the approved protocol, plus a caregiver able to participate. Specify measure, threshold, impairment, time since loss, and any exclusions with the clinical team. A screening cut-off alone is not a diagnosis. Define an appropriate immediate-care route for youth whose needs exceed this program; participation must not delay needed therapy.

Randomize **households**, so siblings do not receive contradictory kit/coaching assignments. Record all participating children and whether a conversation was shared or individual. Before launch, choose either one pre-specified index child for the primary endpoint or a model that includes all eligible children with household clustering. Do not decide after seeing outcomes.

For the first main trial, choose a coherent service population, such as families recruited through bereavement centers receiving the program before specialized therapy, with usual care documented. Study teletherapy adjunct use as a separate cohort/estimand or a sufficiently planned stratum. The adjunct question is whether the program adds benefit to ongoing therapy, not whether a lower-intensity program can substitute for it.

Stratify randomization on a small set of important variables, such as site and broad age band, using a concealed server-side procedure designed by the statistician. Avoid an enormous cross-product of small strata. Loss relationship/circumstance, baseline severity, treatment use, language and socioeconomic/access measures can be planned covariates or exploratory moderators; do not promise powered subgroup conclusions for every loss type.

Children may be nested within households, facilitators, coaching staff or groups. Record those IDs at delivery. Group delivery and shared intervention staff can create dependence even if households are randomized individually. Model it and include it in power simulations; do not treat every child as statistically independent by default. [NIH guidance on individually randomized group-treatment trials](https://researchmethodsresources.nih.gov/methods/irgt).

### Outcomes and a clock that includes non-completers

Select one age-appropriate, licensed/permissioned primary measure of elevated grief-related symptoms with the clinical/research team. Prefer youth self-report for the primary grief construct, with caregiver and functional measures as complementary views. Review language versions, scoring and suitability across the full age range. The existing short screens and in-game thermometers should not automatically become the primary efficacy endpoint.

| Time point | Proposed scheduling anchor | Purpose |
|---|---|---|
| Baseline | Before randomization | Eligibility, symptoms, functioning, context and expectations |
| Midpoint | Day 21 after randomization | Brief mechanism and process assessment; keep burden small |
| Primary post | Day 49 after randomization, pre-specified window | Primary outcome after the planned six-week intervention |
| Maintenance 1 | 90 days after the planned week-6 endpoint | Sustained symptoms/functioning and outside care |
| Maintenance 2 | 180 days after the planned week-6 endpoint | Longer maintenance, renewed distress and service use |

These are proposed anchors and need protocol approval. Crucially, a late game, missed discussion, missed post survey, replacement kit, or treatment discontinuation does not move the research clock. Rescheduled care can have its own dates. Retain research follow-up when participation stops if the family continues to permit it. Missed prior waves never suppress later waves.

Secondary outcomes: functioning/impairment, parent confidence in supporting the child, child-rated comfort in family conversations, parent distress, other clinically relevant symptoms, and use of additional therapy. Mechanisms: whether conversations occurred, felt supportive, and increased understanding or coping confidence. Engagement measures: game/module completion, time, discussion attempts and dose. Costs: kit, shipping, replacements, staff time, supervision and additional contacts. Monitor adverse experiences and clinical escalation under the protocol.

Avoid using completion as a proxy for clinical benefit. More disclosure is not automatically better. Fewer ordinary expressions of missing someone is not necessarily the goal. Parent and child reports can disagree without either being dismissed.

### Analysis and optimization decision

Use intention-to-treat for assigned components; keep allocation separate from actual exposure. Pre-specify the primary endpoint, main-effect estimands, interaction handling, baseline adjustment, clustering, multiplicity and missing-data sensitivity analyses. Report effects and uncertainty, not merely which p-values cross a threshold.

Plan response procedures for missing assessments, document reasons where available, and use methods justified by the missingness assumptions, with sensitivity analyses for informative dropout. Do not analyze completers only or code missed measures as zero symptoms. Follow-up collection should be independent of the coach when possible; assessors should be masked to components where feasible. Family and coach blinding to a kit or phone call is generally impractical.

Pre-specify a decision rule that weighs clinically meaningful benefit, uncertainty, harm, equity and incremental cost. For example: retain components that plausibly improve the pre-defined grief outcome or an agreed mechanism without exceeding the service's cost/capacity ceiling; require a confirmatory package trial before efficacy claims. The clinical team and statistician must define the numerical threshold before outcomes are unmasked. An inconclusive result is not proof that a component has no value.

Outside therapy, emergency support, and unplanned caregiver callbacks must be allowed as needed, logged, and included in interpretation. They are not automatically protocol violations. A per-protocol exposure analysis may supplement, but cannot replace, the randomized comparison. Likewise, an observed association between better conversations and improvement is not by itself proof of mediation.

### Sample-size realism

The number depends on the minimum worthwhile effect, outcome variance, baseline correlation, family/site/staff clustering, attrition, primary comparisons and any interaction claims. Use pilot estimates and sensitivity ranges in simulation rather than a single optimistic effect size from an open trial.

For scale only: a simple two-group main-effect calculation for standardized difference 0.35, 80% power, and two-sided alpha 0.025 gives approximately 311 complete independent outcomes; allowing 20% attrition gives approximately 388 enrollments. That assumes one independent child per household, ignores clustering and baseline precision gains, and is not the final sample size. In a balanced 2 × 2 design, each main contrast uses half the sample against half—not one cell against one cell. Interactions and subgroup claims can require substantially more information. A feasibility study of a few dozen families cannot establish those effects reliably.

## Stage 2: evaluate the selected package

After component selection, run a separate adequately powered randomized trial of the fixed package versus a credible comparison, such as usual bereavement support plus matched general resources. Define usual care and track it rather than treating it as nothing. If the research question is adjunct teletherapy, compare teletherapy plus package with teletherapy under the defined comparison condition.

Use the same independent outcome schedule and maintenance follow-up. Be explicit about the claim supported: an incremental effect of this family program in this setting, not automatic proof of every MGT component, every loss circumstance, or replacement of therapy. Avoid relying on a passive waitlist alone if contact, expectation and access differences would dominate interpretation.

## Later questions worth testing

| Decision | Suitable later design |
|---|---|
| Who needs more coaching after initially struggling? | SMART: pre-specify a meaningful response rule and randomize additional support among eligible nonresponders; clinical rescue is never withheld |
| When should a gentle reminder be offered? | Micro-randomized study of low-risk reminder opportunities, consented channels and quiet hours; measure near-term response and burden |
| Does a booster maintain gains? | A separately planned randomized booster offer at a fixed follow-up point, with outcomes extending beyond it |
| Which kit pieces are worth keeping? | Smaller usability studies, then a targeted component comparison if an important cost/effect uncertainty remains |

These should follow a stable core. A complicated adaptive backend before a clear response rule creates complexity without answering a causal question.

## What the backend must do for the trial

1. Store eligibility and program/research permissions separately; record consent/assent versions and withdrawals.
2. Allocate once per household using an immutable study version; reveal only the component instructions each delivery role needs. Family re-enrollment cannot redraw an allocation.
3. Generate independent assessment dates at allocation and maintain them through missed sessions and surveys.
4. Preserve assignment, delivery, receipt, participation, additional care and deviations as distinct records.
5. Keep study content and templates frozen/versioned; log urgent amendments and which families received each version.
6. Give masked assessors a separate worklist and communication templates; prevent assessment emails from exposing treatment assignments.
7. Export a documented analysis dataset with a data dictionary, stable household/child/site/group/staff codes, denominators, missingness and a reproducible extraction version.
8. Apply role/site/case access on every route; protect addresses, contact details and free text from routine research/fulfillment exports.
9. Support withdrawals by purpose and channel. Record whether permission for previously collected data and future follow-up remains under the approved consent terms.
10. Keep clinical support and urgent response available independently of assignment, and record resolution—not merely notification.

Seek IRB determination/approval for the actual study, including parental permission, child assent, recruitment, observation/recording, compensation and data use. Eligibility for exceptions or waivers is not a product setting. Register prospectively and pre-specify the protocol and analysis plan; confirm funder/journal requirements. NIH's clinical trial registration/reporting policy includes funded behavioral intervention trials. [HHS research-with-children guidance](https://www.hhs.gov/ohrp/regulations-and-policy/guidance/faq/children-research/index.html), [NIH registration and reporting](https://grants.nih.gov/grants/policy/nihgps/HTML5/section_4/4.1.3_clinical_trials_registration_and_reporting_in_clinicaltrials.gov_requirement.htm).

## Decisions to bring to the next planning meeting

1. Is the primary first question delivery optimization or overall clinical effectiveness?
2. Which setting/population is the first target, and how will existing teletherapy be handled?
3. Is the core caregiver task fixed, or is its structured form one of the randomized factors?
4. What clinical improvement and maximum incremental cost would make a component worth retaining?
5. Who performs assessments, owns clinical response, and funds follow-up after treatment stops?
6. How many households can the partners realistically recruit, and how many staff/groups will deliver care?

My default is: finish and observe the six-week family workflow; test kit and proactive coaching; evaluate the selected package; study adaptive support and maintenance boosters later.
