Search the Library
NOTE: This is a new search platform (as of May 2026). If you do a search and don’t get the results you were expecting, please email us at ctnlib@uw.edu to let us know? (If possible, please share your exact search strategy. Thank you!)
Enter keywords and hit Enter (or click the magnifying glass) to search. You can then also select document type or subject/topic to narrow results further (or just use those for searching without a keyword). Results display below this search form.
Document types
Subjects
- CTN-#### format for protocols (CTN-0001, e.g.)
- “exact phrase” (if phrase is not found, it will return results that contain all terms
- word1 NOT word2
- word1 word2 (finds both words)
- Click title to access full-text
- “Show details” reveals abstract & other info
- Checkboxes select items for copy/pasting or printing
- Need help getting a copy of a journal article?
Email ctnlib@uw.edu
Search results
Pragmatic trials provide the opportunity to study the effectiveness of health interventions to improve care in real-world settings. However, use of open-cohort designs with patients becoming eligible after randomization and reliance on electronic health records (EHRs) to identify participants may lead to a form of selection bias referred to as identification bias. This bias can occur when individuals identified as a result of the treatment group assignment are included in analyses.
To demonstrate the importance of identification bias and how it can be addressed, the authors consider a motivating case study, the CTN PRimary care Opioid Use Disorders treatment (PROUD) Trial. PROUD is an ongoing pragmatic, cluster-randomized implementation trial in six health systems to evaluate a program for increasing medication treatment of opioid use disorders (OUDs). A main study objective is to evaluate whether the PROUD intervention decreases acute care utilization among patients with OUD (effectiveness aim). Identification bias is a particular concern, because OUD is underdiagnosed in the EHR at baseline, and because the intervention is expected to increase OUD diagnosis among current patients and attract new patients with OUD to the intervention site. We propose a framework for addressing this source of bias in the statistical design and analysis.
The statistical design sought to balance the competing goals of fully capturing intervention effects and mitigating identification bias, while maximizing power. For the primary analysis of the effectiveness aim, identification bias was avoided by defining the study sample using pre-randomization data (pre-trial modeling demonstrated that the optimal approach was to use individuals with a prior OUD diagnosis). To expand generalizability of study findings, secondary analyses were planned that also included patients newly diagnosed post-randomization, with analytic methods to account for identification bias.
Conclusions: As more studies seek to leverage existing data sources, such as EHRs, to make clinical trials more affordable and generalizable and to apply novel open-cohort study designs, the potential for identification bias is likely to become increasingly common. This case study highlights how this bias can be addressed in the statistical study design and analysis.
Related protocols: CTN-0074
Exercise is a promising treatment for substance use disorders, yet an intention-to-treat analysis of a large, multi-site study found no reduction in stimulant use for exercise versus health education. Exercise adherence was sub-optimal, therefore secondary post-hoc complier average causal effects (CACE) analysis was conducted to determine the potential effectiveness of adequately dosed exercise.
The STimulant use Reduction Intervention using Dosed Exercise (STRIDE) study was a randomized controlled trial comparing a 12 kcal/kg/week (KKW) exercise dose versus a health education control conducted at 9 residential substance use treatment settings across the U.S. that are affiliated with the National Drug Abuse Treatment Clinical Trials Network. Participants were sedentary but medically approved for exercise, used stimulants within 30 days prior to study entry, and received a DSM-IV stimulant abuse or dependence diagnosis within the past year. A CACE analysis adjusted to include only participants with a minimum threshold of adherence (at least 8.3 KKW) and using a negative-binomial hurdle model focused on 218 participants who were 36.2% female, mean age 39.4 years (SD=11.1), and averaged 13 (SD=9.2) stimulant use days in the 30 days before residential treatment. The outcome was days of stimulant use as assessed by the self-report TimeLine Follow Back and urine drug screen results.
The CACE-adjusted analysis found a significantly lower probability of relapse to stimulant use in the exercise group versus the health education group (41% vs. 55.7%, p<.01) and significantly lower days of stimulant use among those who relapsed (5 days vs. 9.9 days, p<.01).
Conclusions: The CACE efficacy analysis was conducted for the STRIDE study to account for exercise dose in the evaluation of exercise as a potential treatment for stimulant use disorders. This analysis demonstrated statistically significant differences for the probability of stimulant use, such that those who would achieve an adequate exercise dose (defined to be an average of 8.3 KKW or more) have an estimated lower probability of relapsing to any stimulant use. Analyses also demonstrated that, even among those who relapsed, the amount of estimated stimulant use was significantly less among those who would achieve an adequate exercise dose. Together, these results suggest a beneficial effect of exercise in the treatment of stimulant abuse. Further research is warranted to develop strategies for exercise adherence that can ensure achievement of an exercise dose sufficient to produce a significant treatment effect.
Related protocols: CTN-0037
In randomized controlled trials (RCTs), a common strategy to increase power to detect a treatment effect is adjustment for baseline covariates. However, adjustment with partly missing covariates, where complete cases are only used, is inefficient. This paper considers different alternatives in trials with discrete-time survival data, where subjects are measured in discrete-time intervals while they may experience an event at any point in time. The results of a Monte Carlo simulation study, as well as a case study of randomized trials in smokers with attention deficit hyperactivity disorder (CTN-0029), indicated that single and multiple imputation methods outperform the other methods and increase precision in estimating the treatment effect. Missing indicator method, which uses a dummy variable in the statistical model to indicate whether the value for that variable is missing and sets the same value to all missing values, is comparable to imputation methods. Nevertheless, the power level to detect the treatment effect based on missing indicator method is marginally lower than the imputation methods, particularly when the missingness depends on the outcome.
In conclusion, complete case analysis is wasteful and drops the power level to a large degree, resulting in an undetectable treatment effect. Also, it can introduce bias in the estimate of the treatment effects when the missingness mechanism depends on the outcome variable. This method is therefore invalid and the authors do not recommend it. Instead, it appears that imputation of partly missing (baseline) covariates should be preferred in the analysis of discrete-time survival data.
Related protocols: CTN-0029
The purpose of this study was to estimate how results would have varied if a substance abuse clinical trial had been conducted with nationally representative adults with substance use and with representative adults receiving substance use treatment. Results were analyzed from NIDA Clinical Trials Network protocol CTN-0044, a multisite clinical trial comparing the effectiveness of the Therapeutic Education System to treatment as usual for outpatient addiction treatment (n = 507). Patients were recruited between June 2010 and August 2011. Abstinence was the primary outcome. The general population sample and general population-treated samples were derived from Wave 1 of the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC) (n = 43,093). Propensity scores provided a standardized measure of the difference between clinical trial participants and the 2 NESARC samples. The clinical trial was reanalyzed by reweighting the sample with propensity scores derived from the 2 samples to obtain generalizable estimates of treatment effects.
Before the clinical trial sample was reweighted, the odds ratio (OR) of response to Therapeutic Education System versus treatment as usual in the trial was 1.62 (95% CI, 1.12-2.35). After the sample was reweighted to be representative of the 2 NESARC groups, ORs were 1.33 (95% CI, 0.34-5.26) for the representative sample with any substance use and 1.64 (95% CI, 0.82-3.27) for the representative treated sample. The effect size of the original study was statistically significant; the estimate effect size for the nationally representative sample was not. This does not necessarily mean that the Therapeutic Education System is not efficacious for the treatment of substance use disorders. Instead, the width of the confidence intervals reflects increased uncertainty associated with extrapolating the results of the clinical trial sample to broader populations.
Conclusions: Applying propensity score weighting to clinical trial results provides a method for estimating the population generalizability of clinical trial findings that relies on effect moderators observed in the study sample and population. Broader confidence intervals in the reweighted samples do not necessarily indicate lack of efficacy of the Therapeutic Education System but rather greater uncertainty concerning effectiveness in general population samples.
Related protocols: CTN-0044
Traditional approaches to subgroup analyses that test each moderating factor as a separate hypothesis can lead to erroneous conclusions due to the problems of multiple comparisons, model misspecification, and multicollinearity. This study aimed to demonstrated a novel, systematic approach to subgroup analyses that avoids these pitfalls. A Best Approximating Model (BAM) approach that identifies multiple moderators and estimates their simultaneous impact on treatment effect sizes was applied to a randomized, controlled, 11-week, double-blind efficacy trial on smoking cessation of adult smokers with attention-deficit/hyperactivity disorder (ADHD), randomized to either OROS-methylphenidate (n=127) or placebo (n=128) and treated with nicotine patch (National Drug Abuse Treatment Clinical Trials Network protocol CTN-0029). Binary outcomes measures were prolonged smoking abstinence and point prevalence smoking abstinence.
Although the original clinical trial data analysis showed no treatment effect on smoking cessation, the BAM analysis showed significant subgroup effects for the primary outcome of prolonged smoking abstinence: (1) lifetime history of substance use disorders, and (2) more severe ADHD symptoms. A significant subgroup effect was also shown for the secondary outcome of point prevalence smoking abstinence — age 18-29 years.
Conclusions: The BAM analysis resulted in different conclusions about subgroup effects compared to a hypothesis-driven approach. These divergent findings underscore the need for investigators to consider more advanced statistical methods to better analyze subgroup effect sizes. By examining moderator independence and avoiding multiple testing, BAMs have the potential to better identify and explain how treatment effects vary across subgroups in heterogeneous patient populations, thus providing better guidance to more effectively match individual patients with specific treatments.
Related protocols: CTN-0029
Recent federal legislation and a renewed focus on integrative care models underscore the need for economical, effective, and science-based behavioral health care treatment. As such, maximizing the impact and reach of treatment research is of great concern. Behavioral health issues, including the frequent co-occurrence of substance use disorders (SUD) and post-traumatic stress disorder (PTSD), are often complex, with a myriad of factors contributing to the success of interventions. Although treatment guides for comorbid SUD/PTSD exist, most patients continue to suffer symptoms following the prescribed treatment course. Further, the study of efficacious treatments has been hampered by methodological challenges (e.g., overreliance on “superiority” designs (i.e., designs structured to test whether or not one treatment statistically surpasses another in terms of effect sizes) and short term interventions). Secondary analyses of randomized controlled clinical trials offer potential benefits to enhance understanding of findings and increase the personalization of treatment.
This paper offers a description of the limits of randomized controlled trials as related to SUD/PTSD populations, highlights the benefits and potential pitfalls of secondary analytic techniques, and uses as a case example one of the largest effectiveness trials of behavioral treatment for co-occurring SUD/PTSD conducted within the National Drug Abuse Treatment Clinical Trials Network (CTN). The paper concludes with implications of this secondary analytic approach to improve addiction researchers’ ability to identify best practices for community-based treatment of these disorders.
Conclusions: Innovative methods are needed to maximize the benefits of clinical studies and better support SUD/PTSD treatment options for both specialty and non-specialty healthcare settings. Given the continuing gap between research and practice, appropriately executed secondary analytic studies are an important step in addressing questions that have real-world value to community clinicians. Moving forward, planning for and description of secondary analyses in randomized trials should be given equal consideration and care to the primary outcome analysis.
This secondary analysis of data from National Drug Abuse Treatment Clinical Trials Network protocol CTN-0003 (“Suboxone (Buprenorphine/Naloxone) Taper: A Comparison of Taper Schedules”) compared three missing data strategies: 1) Latent growth model that assumes the data are missing at random (MAR), 2) Diggle-Kenward missing not at random (MNAR) model where dropout is a function of previous/concurrent urinalysis (UA) submissions, and 3) Wu-Carroll MNAR model where dropout is a function of the growth factors. CTN-0003 examined a 7-day versus 28-day taper for buprenorphine/naloxone to see which taper schedule reduced the likelihood of submitting an opioid-positive UA during treatment.
The MAR model showed a significant effect (B=-0.45, p <0.05) of trial arm on the opioid-positive UA slope (i.e., 28-day taper participants were less likely to submit a positive UA over time) with a small effect size (d=0.20). The MNAR Diggle-Kenward model demonstrated a significant (B=-0.64, p<0.01) effect of trial arm on the slope with a large effect size (d=0.82). The MNAR Wu-Carroll model evidenced a significant (B=-0.41, p<0.05) effect of trial arm on the UA slope that was relatively small (d=0.31).
Conclusions: This performance comparison of three missing data strategies (latent growth model, Diggle-Kenward selection model, Wu-Carrol selection model) on sample data indicates a need for increased use of sensitivity analyses in clinical trial research. Given the potential sensitivity of the trial arm effect to missing data assumptions, it is critical for researchers to consider whether the assumptions associated with each model are defensible.
In randomized controlled trials (RCTs), the most compelling need is to determine whether the treatment condition was more effective than the control. However, it is generally recognized that not all participants in the treatment group of most clinical trials benefit equally. While subgroup analyses are often used to compare treatment effectiveness across pre-determined subgroups categorized by patient characteristics, methods to empirically identify naturally occurring clusters of persons who benefit most from the treatment group have rarely been implemented. This article provides a modeling framework to accomplish this important task.
Utilizing information about individuals from the treatment group who had poor outcomes, the present study proposes an a priori clustering strategy that classifies the individuals with initially good outcomes in the treatment group into: (a) group GE (good outcome, effective), the latent subgroup of individuals for whom the treatment is likely to be effective and (b) group GI (good outcome, ineffective), the latent subgroup of individuals for whom the treatment is not likely to be effective. The method is illustrated through a reanalysis of a publicly available data set from the National Institute on Drug Abuse’s National Drug Abuse Treatment Clinical Trials Network (protocol CTN-0004). That study examined the effectiveness of motivational enhancement therapy from 461 outpatients with substance use disorder problems. As a diagnostic means utilizing out-of-sample forecasting performance, the present study compared the relapse rates during the long-term follow-up period for the two subgroups. As expected, group GI, composed of individuals for whom the treatment was hypothesized to be ineffective, had a significantly higher relapse rate than group GE (63% vs. 27%).
Conclusions: The proposed method, LGEM, identified latent subgroups GE and GI, and the comparison between the two groups revealed several significantly different and informative characteristics even though both subgroups had good outcomes during the immediate post-therapy period. LGEM has potential as a means of further exploring reasons why individuals respond to treatment conditions, regardless of which treatment arm they are exposed to, and can be implemented after the trial is completed, without need for a pre-specified design and can be used by any type of RCT in a variety of topic areas.
Related protocols: CTN-0004
Use of psychosocial measures with different conceptual meanings across cultural groups may render treatment outcome analyses invalid in social work research. Determining measurement invariance allows researchers to assess whether the construct of a measure is similarly comprehended and measured across participant groups (e.g., based on race, ethnicity, gender, age, etc.). Nonequivalence is introduced when groups of participants experience or conceptualize a construct differently, or use distinctive criteria to describe the concept. Measurement nonequivalence across cultural groups is posited to occur due to (a) cultural differences in norms and relevance of the constructs being assessed; (b) language of assessment; or (c) potential differences in participants’ environments and opportunity structures to engage in certain behaviors or develop beliefs due to contextual differences, racism, or other forms of discrimination.
To illustrate this statistical procedure, this poster presents measurement invariance properties across racial groups for two commonly used instruments in social work and substance abuse treatment research (the Revised Helping Alliance Questionnaire (HAq-II) and the Short Inventory of Problems (SIP-R)), using data from the National Drug Abuse Treatment Clinical Trials Network protocol CTN-0004 (“Motivational Enhancement Treatment to Improve Treatment Engagement and Outcome in Subjects Seeking Treatment for Substance Abuse”). Analysis shows that use of measures with different conceptual meaning across racial and ethnic groups may render invalid analyses comparing such groups. Conclusions drawn from invalid findings can lead to ineffective treatments and policy initiatives.
Conclusions: Findings support the comparable understanding of therapeutic alliance and consequences of substance as measured by the HAq-II and SIP-R in African American and non-Latino white participants. Difference in reliability caused by the identified items needs verification in future studies to ensure use of reliable HAq-II and SIP-R latent factors.
Related protocols: CTN-0004
In case multiple treatment alternatives are available for some medical problem, the detection of treatment-subgroup interactions (i.e., relative treatment effectiveness varying over subgroups of persons) is of key importance for personalized medicine and the development of optimal treatment strategies. Randomized clinical trials (RCTs) often go without clear a priori hypotheses on the subgroups involved in treatment-subgroup interactions, and with a large number of pre-treatment characteristics in the data. In such situations, relevant subgroups (defined in terms of pre-treatment characteristics) are to be induced during the actual data analysis. This comes down to a problem of cluster analysis, with the goal of this analysis being to find clusters of persons that are involved in meaningful treatment-person cluster interactions. For such a cluster analysis, five recently proposed methods can be used, all being of a recursive partitioning type. However, these five methods have been developed almost independently, and the relations between them are not yet understood.
This paper aims to close that gap. It starts by outlining the basic principles behind each method, and by illustrating it with an application on a data set from an RCT in the National Drug Abuse Treatment Clinical Trials Network that evaluated two treatment strategies for substance abuse problems (CTN-0005, “Motivational Interviewing (MI) to Improve Treatment Engagement and Outcome in Subjects Seeking Treatment for Substance Abuse”). Next, it presents a comparison of the methods, focusing on major similarities and differences. The discussion concludes with practical advice for end users with regard to the selection of a suitable method, and with an important challenge for future research in this area.
Related protocols: CTN-0005
This two-hour webinar, produced by the National Drug Abuse Treatment Clinical Trials Network (CTN) Clinical Coordinating Center for CTN members and the public, features a plain-English description of the intuition behind basic statistical concepts used in clinical trials. Its content requires no statistical background and aims to bridge the communication gap between researchers and biostatisticians. It does not teach how to perform statistical tasks; there are no formulas and no proofs. Instead, it explains why these statistical tasks are performed and what they mean once they are performed.
This webinar is intended for CTN members and the public, especially non-statisticians with experience in clinical trials who seek a better understanding of statistical concepts encountered throughout the cycle of a clinical trial.
Note: The webinar included the showing of two short videos, which you can only hear the sound for in this recording (the picture is blank). If you wish to view the short films, they can both be found on YouTube:
First clip:
http://www.youtube.com/watch?v=AukrpAoAYY0
Second clip:
http://www.youtube.com/watch?v=AukrpAoAYY0
Presented by Paul G. Wakim, PhD (NIDA Center for the Clinical Trials Network) and Abigail G. Matthews, PhD (CTN Data & Statistics Center, EMMES).
Two common procedures for the treatment of missing information, listwise deletion and positive urine analysis (UA) imputation (e.g., if the participant fails to provide urine for analysis, then score the UA positive), may result in significant biases during the interpretation of treatment effects. To compare these approaches and to offer a possible alternative, these two procedures were compared to the multiple imputation (MI) procedure with publicly available data from a recent clinical trial (National Drug Abuse Treatment Clinical Trials Network protocol CTN-0003, Ling et al, 2009). Listwise deletion, single imputation (i.e., positive UA imputation), and MI missing data procedures were used to comparatively examine the effect of the protocol’s two different buprenorphine/naloxone tapering schedules (7- or 28-days) for opioid addiction on the likelihood of a positive UA. The listwise deletion of missing data resulted in a nonsignificant effect for the taper while the positive UA imputation procedure resulted in a significant effects, replicating the original findings by Ling et al (2009). Although the MI procedure also resulted in a significant effect, the effect size was meaningfully smaller and the standard errors meaningfully larger when compared to the positive UA procedure. This study demonstrates that the researcher can obtain markedly different results depending on how the missing data are handled. Missing data theory suggests that listwise deletion and single imputation procedures should not be used to account for missing information, and that MI has advantages with respect to internal and external validity when the assumption of missing at random can be reasonably supported. Consistent with previous investigation of missing data in substance abuse treatment, the authors encourage researchers to understand and report the missing data mechanism as well as use newer procedures for the treatment of missing information (i.e., MI or direct maximum likelihood procedures) that are based on a research-specified “best estimate” of the missing values.
Related protocols: CTN-0003
Selection of appropriate outcome measures is important for clinical studies of drug addiction treatment. Researchers use various methods for collecting drug use outcomes and must consider substances to be included in a urine drug screen (UDS), accuracy of self-report, use of various instruments and procedures for collecting self-reported drug use, and timing of outcome assessments. This study sought to define a set of candidate measures to (1) assess their intercorrelation and (2) identify any differences in results. To that end, data were combined from seven completed protocols in the National Drug Abuse Treatment Clinical Trials Network (CTN), with a total of 1897 participants. Nine outcome measures were defined, based on UDS, self-report, or a combination, then multivariable, multilevel generalized estimating equation models were used to assess subgroup differences in intervention success, controlling for baseline differences and accounting for clustering by CTN protocols. Results found high correlations among all candidate outcomes. All outcomes showed consistent overall results with no significant intervention impact on drug use during follow-up. However, with most UDS variables, but not with self-report or “corrected self-report,” a significant gender–ethnicity interaction with benefit shown in African American women, White women, and Hispanic men was observed.
Conclusions: Despite strong associations between candidate measures, important differences in results were found. This study demonstrates the potential utility and impact of combining UDS and self-report data for drug use assessment. The results suggest possible differences in intervention efficacy by gender and ethnicity, but highlight the need to cautiously interpret observed interactions. Additional studies like this one will help guide implementation of methodological recommendations to construct combined measures.
Clinical trials testing the effectiveness of interventions for addictions, HIV transmission risk, and other behavioral health problems are important to advancing evidence-based treatment. Such trials are expensive and time-consuming to conduct, but the underlying effect sizes tend to be modest, and often findings are disappointing, failing to show evidence of treatment effects. This study aimed to demonstrate how appropriate covariation for baseline severity can enhance detection of treatment effects, using an example from the National Drug Abuse Treatment Clinical Trials Network (protocol CTN-0015, “Women and Trauma”). Baseline severity, the score of the outcome measure at baseline, prior to randomization, is often strongly associated with outcome in such studies. Covariation for baseline score may enhance detection of treatment effects, because the variance explained by the baseline score is removed from the error variance in the estimate of the difference in outcome between treatments. Alternatively, the effect of treatment may manifest in the form of a baseline-by-treatment interaction. Common interaction patterns include that treatment may be more effective among patients with higher levels of baseline severity, or treatment may be more effective among patients with low severity at baseline (“relapse prevention” effect). Such effects may be important to developing treatment guidelines and offer clues toward understanding the mechanisms of action of treatments and of the disorders.
Conclusions: This article illustrates principles of covariation for baseline and the baseline-by-treatment interaction in nontechnical graphical terms, and discusses examples from clinical trials, including the CTN. Implications for the design and analysis of clinical trials are discussed, and it is argued that covariation for baseline severity of the outcome measure and testing of the baseline-by-treatment interaction should be considered for inclusion in the primary outcome analyses of treatment effectiveness trials of substantial size.
In clinical trials of treatment for stimulant abuse, including several National Drug Abuse Treatment Clinical Trials Network (CTN) protocols, researchers commonly record both Time-Line Follow-Back (TLFB) self-reports and urine drug screen (UDS) results. This study aimed to compare the power of self-report, qualitative (use vs. no use) UDS assessment, and various algorithms to generate self-report-UDS composite measures to detect treatment differences via t-test in simulated clinical trial data. Monte Carlo simulations, patterned in part on real data to model self-report reliability, were performed on UDS errors, dropout, informatively missing UDS reports, incomplete adherence to a urine donation schedule, temporal correlation of drug use, number of days in the study period, number of patients per arm, and distribution of drug-use probabilities. Investigated algorithms include maximum likelihood and Bayesian estimates, self-report alone, UDS alone, and several simple modifications of self-report (referred to here as ELCON algorithms) which eliminate perceived contradictions between it and UDS. Among the algorithms investigated, simple ELCON algorithms gave rise to the most powerful t-tests to detect mean group differences in stimulant drug use.
Conclusions: Further investigation is needed to determine if simple, naïve procedures such as the ELCON algorithms are optimal for comparing clinical study treatment arms. But researchers who currently require an automated algorithm in scenarios similar to those simulated for combining TLFB and UDS to test group differences in stimulant use should consider one of the ELCON algorithms. This analysis continues a line of inquiry which could determine how best to measure outpatient stimulant use in clinical trials.