<!DOCTYPE article
PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with MathML3 v1.4 20241031//EN" "JATS-archivearticle1-4-mathml3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="1.4" xml:lang="en" article-type="article-commentary"><?properties manuscript?><processing-meta base-tagset="archiving" mathml-version="3.0" table-model="xhtml" tagset-family="jats"><restricted-by>pmc</restricted-by></processing-meta><front><journal-meta><journal-id journal-id-type="nlm-journal-id">7802871</journal-id><journal-id journal-id-type="pubmed-jr-id">4294</journal-id><journal-id journal-id-type="nlm-ta">Int J Epidemiol</journal-id><journal-id journal-id-type="iso-abbrev">Int J Epidemiol</journal-id><journal-title-group><journal-title>International journal of epidemiology</journal-title></journal-title-group><issn pub-type="ppub">0300-5771</issn><issn pub-type="epub">1464-3685</issn></journal-meta><article-meta><article-id pub-id-type="pmid">41370626</article-id><article-id pub-id-type="pmc">12694401</article-id><article-id pub-id-type="doi">10.1093/ije/dyaf207</article-id><article-id pub-id-type="manuscript">NIHMS2192435</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title-group><article-title>Making DAGs Even More Useful: Using Augmented Causal Diagrams to Depict Counterfactual, Study Design, Measurement, Analytical, and Interventional Features</article-title></title-group><contrib-group><contrib contrib-type="author"><name><surname>Arah</surname><given-names>Onyebuchi A.</given-names></name><xref rid="A1" ref-type="aff">1</xref><xref rid="A2" ref-type="aff">2</xref><xref rid="A3" ref-type="aff">3</xref><xref rid="A4" ref-type="aff">4</xref><xref rid="CR1" ref-type="corresp">*</xref></contrib><contrib contrib-type="author"><name><surname>Coates</surname><given-names>Matthew M.</given-names></name><xref rid="A1" ref-type="aff">1</xref><xref rid="A2" ref-type="aff">2</xref></contrib></contrib-group><aff id="A1"><label>1</label>Department of Epidemiology, Fielding School of Public Health, University of California, Los Angeles (UCLA), Los Angeles, CA, USA</aff><aff id="A2"><label>2</label>Practical Causal Inference Lab, UCLA, Los Angeles, CA, USA</aff><aff id="A3"><label>3</label>Department of Statistics &#x00026; Data Science, UCLA, Los Angeles, CA, USA</aff><aff id="A4"><label>4</label>Research Unit for Epidemiology, Department of Public Health, Aarhus University, Aarhus, Denmark</aff><author-notes><fn fn-type="con" id="FN1"><p id="P1">Author contributions</p><p id="P2">OAA drafted the commentary and will act as its guarantor. OAA and MMC contributed to this work and approved the final submission.</p></fn><corresp id="CR1"><label>*</label>Corresponding author. Department of Epidemiology, UCLA Fielding School of Public Health, 650 Charles E. Young Drive South, Los Angeles, CA 90095, USA. <email>arah@ucla.edu</email></corresp></author-notes><pub-date pub-type="nihms-submitted"><day>3</day><month>7</month><year>2026</year></pub-date><pub-date pub-type="ppub"><day>14</day><month>10</month><year>2025</year></pub-date><pub-date pub-type="pmc-release"><day>17</day><month>8</month><year>2026</year></pub-date><volume>54</volume><issue>6</issue><elocation-id>dyaf207</elocation-id><kwd-group><kwd>directed acyclic graph</kwd><kwd>augmented directed acyclic graph</kwd><kwd>causal diagram</kwd><kwd>graphical augmentation</kwd><kwd>disease risk score</kwd><kwd>propensity score</kwd></kwd-group></article-meta></front><body><sec id="S1"><title>Introduction</title><p id="P3">Since their mainstream introduction in the 1990s, causal diagrams, including directed acyclic graphs (DAGs), are increasingly used to depict our causal knowledge of the world we study and to guide study design, analysis, and interpretation for causal inference [<xref rid="R1" ref-type="bibr">1</xref>, <xref rid="R2" ref-type="bibr">2</xref>]. Beyond describing the data-generating mechanism, DAGs are typically used to select variables for confounding control [<xref rid="R2" ref-type="bibr">2</xref>]. However, to do more, researchers have had to modify or augment their DAGs with additional features reflecting theoretical what-if scenarios or study-dependent processes affecting the data, analysis, and interpretation. In this journal, Mansournia <italic toggle="yes">et al</italic>. [<xref rid="R3" ref-type="bibr">3</xref>] show how to depict balancing scores on DAGs, a welcome addition to the growing use of augmented graphs. This commentary will first place this work in the broader context of how researchers have tried to do more with augmented causal diagrams, including augmented directed acyclic graphs (ADAGs) which we use more expansively to include all such graphs in this commentary. Then, we highlight some useful features of ADAGs that depict balancing scores, focusing on the propensity score for illustration. We conclude with some remarks on the future of graphical augmentation.</p></sec><sec id="S2"><title>Making DAGs even more useful with augmentation</title><p id="P4">Over the years, DAGs have been augmented to depict counterfactual variables and study design, measurement, analytical, and interventional features [<xref rid="R1" ref-type="bibr">1</xref>, <xref rid="R4" ref-type="bibr">4</xref>&#x02013;<xref rid="R10" ref-type="bibr">10</xref>], yielding what we call augmented DAGs (ADAGs). Consider the nonaugmented DAG in <xref rid="F1" ref-type="fig">Figure 1a</xref> with an exposure <italic toggle="yes">A</italic>, an outcome <italic toggle="yes">Y</italic>, a mediator <italic toggle="yes">M</italic>, and two confounders <italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2, which could be used to study the effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic> and the joint effect of <italic toggle="yes">A</italic> and <italic toggle="yes">L</italic>1 and <italic toggle="yes">Y</italic>. Suppose we are interested in visualizing non-differential independent measurement error in <italic toggle="yes">A</italic>, effect-measurement modification of the effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic> by <italic toggle="yes">L</italic>1, interaction between <italic toggle="yes">A</italic> and <italic toggle="yes">M</italic> in causing <italic toggle="yes">Y</italic>, that the confounder <italic toggle="yes">L</italic>2 is unmeasured but its proxies <italic toggle="yes">E</italic>2 and <italic toggle="yes">O</italic>2 are observed, and the consequences of selection on unmeasured <italic toggle="yes">L</italic>2 which affects whether <italic toggle="yes">Y</italic> is observed.</p><p id="P5">These study-dependent processes or features can be added to DAG 1a to obtain the ADAG in <xref rid="F1" ref-type="fig">Figure 1b</xref>. For example, the measurement error in exposure is added as a new node <italic toggle="yes">A</italic>* with an arrow from <italic toggle="yes">A</italic> [<xref rid="R4" ref-type="bibr">4</xref>, <xref rid="R5" ref-type="bibr">5</xref>]. The effect-measure modification of <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">Y</italic> by <italic toggle="yes">L</italic>1 is visualized using a product node <italic toggle="yes">AL</italic>1 with arrows from <italic toggle="yes">A</italic> and <italic toggle="yes">L</italic>1 and an outgoing arrow into <italic toggle="yes">Y</italic> [<xref rid="R6" ref-type="bibr">6</xref>]. Similar augmentation draws the interaction between <italic toggle="yes">A</italic> and <italic toggle="yes">M</italic> in affecting <italic toggle="yes">Y</italic>. The selection mechanism that results from selection on <italic toggle="yes">S</italic> = 1 and how <italic toggle="yes">S</italic> defines whose <italic toggle="yes">Y</italic> is observed in the study are shown, respectively, with an arrow from unobserved <italic toggle="yes">L</italic>2 (a parent of <italic toggle="yes">Y</italic>) into <italic toggle="yes">S</italic> and arrows from <italic toggle="yes">S</italic> and unobserved <italic toggle="yes">Y</italic> into observed <italic toggle="yes">Y</italic><sup>obs</sup> in the ADAG 1b [<xref rid="R7" ref-type="bibr">7</xref>, <xref rid="R8" ref-type="bibr">8</xref>]. A signed ADAG depicting the direction of effects with + and &#x02013; symbols can be used for reasoning about uncontrolled confounding (due to <italic toggle="yes">L</italic>2 here) (<xref rid="F1" ref-type="fig">Figure 1b</xref>). Others, such as single-world intervention graphs (as in ADAG 1c) and twin networks (as in ADAG 1d), have been developed to show interventions and thus potential outcomes [chapter 41 in reference <xref rid="R1" ref-type="bibr">1</xref>]. Although standard d-separation can be used to read some key independencies (e.g., conditional exchangeability) for using the PS in these ADAGs to estimate <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">Y</italic>, other independence implications of deterministic augmentations need the (modified) D-separation criterion [<xref rid="R9" ref-type="bibr">9</xref>].</p></sec><sec id="S3"><title>Doing more with ADAGs that depict balancing scores</title><p id="P6">Given the growing use of ADAGs, it is unsurprising to see ones depicting balancing scores, namely propensity scores (PS) and disease risk scores (DRS), used in confounding control [<xref rid="R3" ref-type="bibr">3</xref>]. Like others [<xref rid="R10" ref-type="bibr">10</xref>], Mansournia <italic toggle="yes">et al</italic>. [<xref rid="R3" ref-type="bibr">3</xref>] propose inserting the PS between the exposure <italic toggle="yes">A</italic> and the confounders <italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2, and the DRS between the confounders and the (potential) outcome. [By the way, we believe Mansournia <italic toggle="yes">et al</italic>&#x02019;s arrow from <italic toggle="yes">A</italic> to the computational PS in their Figure 8c should be omitted.] Is it possible to anticipate or design the theoretical balancing scores from the unmodified DAG, like the one in <xref rid="F1" ref-type="fig">Figure 1e</xref>? Is there a general principle for augmenting DAGs with the balancing scores? Which variables should be included? Can we evaluate variable balance using the resulting ADAGs? What else can we infer from them?</p><p id="P7">Consider the DAG in <xref rid="F1" ref-type="fig">Figure 1e</xref>, adapted from Mansournia <italic toggle="yes">et al</italic>. by including <italic toggle="yes">L</italic>0 (which affects exposure <italic toggle="yes">A</italic> directly and <italic toggle="yes">Y</italic> only through <italic toggle="yes">A</italic>), <italic toggle="yes">L</italic>4 (which affects <italic toggle="yes">Y</italic> only), and a collider <italic toggle="yes">L</italic>3 (which forms an <italic toggle="yes">M</italic>-bias structure between <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic>). Under Pearl&#x02019;s Structural Causal Model [<xref rid="R1" ref-type="bibr">1</xref>] that would have given rise to DAG 1e, <italic toggle="yes">A</italic> is given by the function <italic toggle="yes">f</italic><sub>A</sub>(<italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">U</italic><sub>1</sub>, <italic toggle="yes">U</italic><sub>A</sub>) with <italic toggle="yes">U</italic><sub>A</sub> reflecting the unknown parents of A that are independent of the unknown parents of the other variables in DAG 1e. Then, imagine inventing the PS from DAG 1e by partitioning <italic toggle="yes">A</italic> into two independent parts which would have arrows from them into <italic toggle="yes">A</italic>: a deterministic part we call the PS which is formed by the <italic toggle="yes">L</italic>s in <italic toggle="yes">f</italic><sub>A</sub> (and represented with arrows from the <italic toggle="yes">L</italic>s to the PS in ADAGs 1f to 1j), and a variable &#x02018;error&#x02019; part containing at least <italic toggle="yes">U</italic><sub>A</sub>. Therefore, a general principle for representing the PS as a function of selected <italic toggle="yes">L</italic>s on the ADAG involves moving any existing d-connection between the <italic toggle="yes">L</italic>s and <italic toggle="yes">A</italic> from <italic toggle="yes">L</italic>s to PS and an arrow from PS to <italic toggle="yes">A,</italic> as shown in ADAGs 1f to 1j. Nodes like <italic toggle="yes">L</italic>3 or <italic toggle="yes">L</italic>4, which do not cause <italic toggle="yes">A</italic> but are used to specify PS, can be connected using dashed arrows to distinguish their non-causal role in the partitioning of <italic toggle="yes">A</italic>. We could conduct a similar visual thought experiment to &#x02018;invent&#x02019; the DRS on an ADAG using the structural model for <italic toggle="yes">Y</italic>, <italic toggle="yes">f</italic><sub>Y</sub>(<italic toggle="yes">A</italic>, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4, <italic toggle="yes">U</italic><sub>2</sub>, <italic toggle="yes">U</italic><sub>Y</sub>).</p><p id="P8">If the <italic toggle="yes">L</italic>s used in forming the PS suffice to close the backdoor between <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic>, then the PS could also close the same backdoor (see ADAGs 1f to 1j). This answers the question of which variables to include in the PS: at a minimum, those <italic toggle="yes">L</italic>s (<italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2) that would suffice to &#x02018;deconfound&#x02019; <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">Y</italic>. Mansournia <italic toggle="yes">et al</italic>. use a SWIG (a causal diagram augmented with interventional features) to illustrate how conditioning on the DRS satisfies the traditional conditional ignorability/exchangeability assumptions, which could also be shown for the PS. <xref rid="T1" ref-type="table">Table 1</xref> presents five PS models using various combinations of <italic toggle="yes">L</italic>0 through <italic toggle="yes">L</italic>4, corresponding to ADAGS 1f to 1j. Only the first four PS models that exclude the collider L3 can be expected to result in no open backdoor between <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic>. From <xref rid="T1" ref-type="table">Table 1</xref>, we also see that a PS model can only be used to achieve variable balance (or independence) between any <italic toggle="yes">L</italic> and <italic toggle="yes">A</italic> if that <italic toggle="yes">L</italic> is included in the PS. Any <italic toggle="yes">L</italic> not included in a PS (e.g., <italic toggle="yes">L</italic>0 or <italic toggle="yes">L</italic>4 for PS1 in ADAG 1f) retains its original (im)balance across levels of <italic toggle="yes">A</italic>. Introducing a collider <italic toggle="yes">L</italic>3 into a PS induces collider-bias (ADAG 1j and PS5 in <xref rid="T1" ref-type="table">Table 1</xref>). ADAGs also give intuition about implications of variable selection for statistical efficiency based on the relationship of selected variables to the exposure and outcome (<xref rid="T1" ref-type="table">Table 1</xref>).</p><p id="P9">Finally, from the ADAGs, we can also infer that we could regress <italic toggle="yes">Y</italic> on the residuals (<italic toggle="yes">U</italic><sub><italic toggle="yes">A</italic></sub>) from the exposure model (<italic toggle="yes">A</italic> = PS + <italic toggle="yes">U</italic><sub><italic toggle="yes">A</italic></sub>) used to create the PS to get unconfounded causal estimates. For example, by using <italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2 to specify PS1 (in ADAG 1f), the error term would contain only information from the variables that are instruments for <italic toggle="yes">A</italic> (i.e. <italic toggle="yes">L</italic>0, <italic toggle="yes">U</italic><sub><italic toggle="yes">A</italic>1</sub>, and the unspecified unmeasured common cause(s) of <italic toggle="yes">A</italic> and <italic toggle="yes">L</italic>3 represented by the dashed bidirected arc).</p></sec><sec id="S4"><title>Conclusion</title><p id="P10">Balancing scores, such as PS and DRS, can be depicted on ADAGs [<xref rid="R3" ref-type="bibr">3</xref>, <xref rid="R10" ref-type="bibr">10</xref>]. Due to space constraints, we focused on PS in DAGs, although similar results can be developed for the DRS. Drawing ADAGs should be used after drawing and using the unmodified DAG for identification. In this commentary, we highlighted only some implications of PS-augmented graphs, but we hope more work will be done on extending them. We argue that ADAGs can be more useful than plain DAGs in empirical causal inference work and deserve further research to help us understand what we can and cannot do with them.</p></sec></body><back><ack id="S5"><title>Funding</title><p id="P11">OAA was supported in part by the Centers for Disease Control and Prevention (CDC) grant agreement number T42OH008412, the National Center for Advancing Translational Sciences (NCATS) of the National Institutes of Health (NIH) under the UCLA Clinical and Translational Science Institute grant number UL1TR001881, and the Karen Toffler Charitable Trust. The National Cancer Institute (NCI/NIH) grant number T32CA009142 supported MMC. Funding does not imply endorsement by or the views of the funders.</p></ack><fn-group><fn id="FN2"><p id="P12">Ethics approval</p><p id="P13">None needed for this commentary since no data were analyzed.</p></fn><fn id="FN3"><p id="P14">Use of artificial intelligence (AI) tools</p><p id="P15">No AI tools were used in or for this commentary.</p></fn></fn-group><ref-list><title>References</title><ref id="R1"><label>1.</label><mixed-citation publication-type="book"><name><surname>Geffner</surname><given-names>H</given-names></name>, <name><surname>Dechter</surname></name>, <name><surname>Halpern</surname><given-names>JY</given-names></name> (Editors). <source>Probabilistic and Causal Inference: The Works of Judea Pearl</source>. <publisher-name>Association for Computing Machinery</publisher-name>, <year>2022</year>.</mixed-citation></ref><ref id="R2"><label>2.</label><mixed-citation publication-type="journal"><name><surname>Tennant</surname><given-names>PWG</given-names></name>, <name><surname>Murray</surname><given-names>EJ</given-names></name>, <name><surname>Arnold</surname><given-names>KF</given-names></name>, <name><surname>Berrie</surname><given-names>L</given-names></name>, <name><surname>Fox</surname><given-names>MP</given-names></name>, <name><surname>Gadd</surname><given-names>SC</given-names></name>, <name><surname>Harrison</surname><given-names>WJ</given-names></name>, <name><surname>Keeble</surname><given-names>C</given-names></name>, <name><surname>Ranker</surname><given-names>LR</given-names></name>, <name><surname>Textor</surname><given-names>J</given-names></name>, <name><surname>Tomova</surname><given-names>GD</given-names></name>, <name><surname>Gilthorpe</surname><given-names>MS</given-names></name>, <name><surname>Ellison</surname><given-names>GTH</given-names></name>. <article-title>Use of directed acyclic graphs (DAGs) to identify confounders in applied health research: review and recommendations</article-title>. <source>Int J Epidemiol</source>
<year>2021</year>;<volume>50</volume>:<fpage>620</fpage>&#x02013;<lpage>632</lpage>.<pub-id pub-id-type="pmid">33330936</pub-id>
</mixed-citation></ref><ref id="R3"><label>3.</label><mixed-citation publication-type="journal"><name><surname>Mansournia</surname><given-names>M</given-names></name>, <name><surname>Nazemipour</surname><given-names>M</given-names></name>, <name><surname>Etminan</surname><given-names>M</given-names></name>. <article-title>Balancing scores and causal diagrams</article-title>. <source>Int J Epidemiol</source>
<year>2025</year>;<volume>54</volume>:<fpage>dyaf114</fpage>.<pub-id pub-id-type="pmid">40623152</pub-id>
</mixed-citation></ref><ref id="R4"><label>4.</label><mixed-citation publication-type="journal"><name><surname>Robins</surname><given-names>JM</given-names></name>. <article-title>Data, design, and background knowledge in etiologic inference</article-title>. <source>Epidemiology</source>
<year>2001</year>;<volume>12</volume>:<fpage>313</fpage>&#x02013;<lpage>320</lpage>.<pub-id pub-id-type="pmid">11338312</pub-id>
</mixed-citation></ref><ref id="R5"><label>5.</label><mixed-citation publication-type="journal"><name><surname>Wardle</surname><given-names>MT</given-names></name>, <name><surname>Reavis</surname><given-names>KM</given-names></name>, <name><surname>Snowden</surname><given-names>JM</given-names></name>. <article-title>Measurement error and information bias in causal diagrams: mapping epidemiological concepts and graphical structures</article-title>. <source>Int J Epidemiol</source>
<year>2025</year>;<volume>53</volume>:<fpage>dyae141</fpage>.</mixed-citation></ref><ref id="R6"><label>6.</label><mixed-citation publication-type="confproc"><name><surname>Arah</surname><given-names>OA</given-names></name>. <article-title>Augmenting causal diagrams with effect modification, interaction, and other parametric information</article-title>. <conf-name>Abstract #283 in the Abstract Book of the 48th Annual Meeting of the Society for Epidemiologic Research</conf-name>, <conf-date>June 16&#x02013;19, 2015</conf-date>. <ext-link xlink:href="https://github.com/oacarah/causal-diagrams.git" ext-link-type="uri">https://github.com/oacarah/causal-diagrams.git</ext-link> (<date-in-citation>9 October 2025</date-in-citation>, date last accessed).</mixed-citation></ref><ref id="R7"><label>7.</label><mixed-citation publication-type="journal"><name><surname>Hern&#x000e1;n</surname><given-names>MA</given-names></name>, <name><surname>Hern&#x000e1;ndez-D&#x000ed;az</surname><given-names>S</given-names></name>, <name><surname>Robins</surname><given-names>JM</given-names></name>. <article-title>A structural approach to selection bias</article-title>. <source>Epidemiology</source>
<year>2004</year>;<volume>15</volume>:<fpage>615</fpage>&#x02013;<lpage>25</lpage>.<pub-id pub-id-type="pmid">15308962</pub-id>
</mixed-citation></ref><ref id="R8"><label>8.</label><mixed-citation publication-type="journal"><name><surname>Arah</surname><given-names>OA</given-names></name>. <article-title>Analyzing selection bias for credible causal inference: When in doubt, DAG it out</article-title>. <source>Epidemiology</source>
<year>2019</year>;<volume>30</volume>:<fpage>517</fpage>&#x02013;<lpage>520</lpage>.<pub-id pub-id-type="pmid">31033691</pub-id>
</mixed-citation></ref><ref id="R9"><label>9.</label><mixed-citation publication-type="journal"><name><surname>Berrie</surname><given-names>L</given-names></name>, <name><surname>Arnold</surname><given-names>KF</given-names></name>, <name><surname>Tomova</surname><given-names>GD</given-names></name>, <name><surname>Gilthorpe</surname><given-names>MS</given-names></name>, <name><surname>Tennant</surname><given-names>PWG</given-names></name>. <article-title>Depicting deterministic variables withing directed acyclic variables: an aid for identifying and interpreting causal effects involving derived variables</article-title>. <source>Am J Epidemiol</source>
<year>2025</year>;<volume>194</volume>:<fpage>469</fpage>&#x02013;<lpage>479</lpage>.<pub-id pub-id-type="pmid">38918044</pub-id>
</mixed-citation></ref><ref id="R10"><label>10.</label><mixed-citation publication-type="journal"><name><surname>Arah</surname><given-names>OA</given-names></name>. <article-title>Using causal diagrams optionally augmented with functional mapping to understand covariate balance in propensity score models. Abstract #445 in the Abstracts of the 42nd Annual Meeting of the Society for Epidemiologic Research, June 23&#x02013;26, 2009</article-title>. <source>Am J Epidemiol</source>
<year>2009</year>;<volume>169</volume>(<issue>suppl 11</issue>):<fpage>S112</fpage>. <ext-link xlink:href="https://github.com/oacarah/causal-diagrams.git" ext-link-type="uri">https://github.com/oacarah/causal-diagrams.git</ext-link> (<date-in-citation>9 October 2025</date-in-citation>, date last accessed).</mixed-citation></ref></ref-list></back><floats-group><fig position="float" id="F1"><label>Figure 1:</label><caption><p id="P16">Nonaugmented and augmented causal diagrams or directed acyclic graphs</p><p id="P17">(a) Nonaugmented directed acyclic graph (DAG) for the effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic>. (b) Augmented DAG (ADAG) showing: (i) non-differential independent measurement error in <italic toggle="yes">A</italic> with node <italic toggle="yes">A</italic>*, (ii) effect measure modification by <italic toggle="yes">L</italic>1 with AL1 as a deterministic product node, (iii) interaction between <italic toggle="yes">A</italic> and <italic toggle="yes">M</italic> in causing <italic toggle="yes">Y</italic> with a similar product node augmentation, (iv) unobserved nature of <italic toggle="yes">L</italic>2 with parentheses but observable proxies <italic toggle="yes">E</italic>2 and <italic toggle="yes">O</italic>2 which can be used for proximal causal inference, (v) a positive sign on arrow from unmeasured confounder <italic toggle="yes">L</italic>2 to <italic toggle="yes">A</italic> and a negative sign on the arrow from <italic toggle="yes">L</italic>2 to <italic toggle="yes">Y</italic> indicating the direction of those links and hence of uncontrolled confounding under some monotonicity assumptions, and (vi) selection of observations in which <italic toggle="yes">Y</italic> is only observed (<italic toggle="yes">Y</italic><sup>obs</sup>) whenever the selection node <italic toggle="yes">S</italic> = 1, and selection is a consequence of an unobserved parent (<italic toggle="yes">L</italic>2) of <italic toggle="yes">Y</italic>. Furthermore, in ADAG 1b, adding the <italic toggle="yes">AM</italic> node allows us to see how the total effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic> can be decomposed into four paths or component effects with respect to mediation and interaction: the direct <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">Y</italic> path (controlled direct effect), <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">AM</italic>&#x02192;<italic toggle="yes">Y</italic> path not through <italic toggle="yes">M</italic> (reference interaction effect), <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">M</italic>&#x02192;<italic toggle="yes">AM</italic>&#x02192;<italic toggle="yes">Y</italic> path (mediated interaction effect), and <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">M</italic>&#x02192;<italic toggle="yes">Y</italic> path not through <italic toggle="yes">AM</italic> (pure indirect effect). An additional 5th path <italic toggle="yes">A</italic>&#x02192;<italic toggle="yes">AL</italic>1&#x02192;<italic toggle="yes">Y</italic> represents a potential <italic toggle="yes">L</italic>1-modified direct effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic>. These component direct paths, although implied by DAG 1a, are not apparent without augmentation. (c) Corresponding single-world intervention graph (SWIG) with some observed nodes and consequences of joint interventions setting <italic toggle="yes">L</italic>1 = <italic toggle="yes">l</italic>1 and <italic toggle="yes">A</italic> = <italic toggle="yes">a</italic>. (d) Corresponding twin network graph with all observed nodes and consequences of intervention setting L1 = <italic toggle="yes">l</italic>1 and <italic toggle="yes">A</italic> = <italic toggle="yes">a</italic>, with shared parents between networks depicted with red arrows. (e) DAG that could be used to study the effect of exposure <italic toggle="yes">A</italic> on outcome <italic toggle="yes">Y,</italic> showing parents of <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic> (<italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2), parent of <italic toggle="yes">A</italic> (<italic toggle="yes">L</italic>0), parent of <italic toggle="yes">Y</italic> (<italic toggle="yes">L</italic>4), a collider (<italic toggle="yes">L</italic>3), and the unknown parents (<italic toggle="yes">U</italic><sub><italic toggle="yes">A</italic></sub>) of <italic toggle="yes">A</italic> being independent of the unknown parents of other variables. (f) ADAG that includes a propensity score PS1 for A specified as a function of <italic toggle="yes">L</italic>1 and <italic toggle="yes">L</italic>2, shown as a deterministic node with a double box. (g) ADAG that includes a propensity score PS2 for <italic toggle="yes">A</italic> specified using <italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, and <italic toggle="yes">L</italic>2. (h) ADAG that includes a propensity score PS3 for <italic toggle="yes">A</italic> specified using <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, and <italic toggle="yes">L</italic>4, with a dashed arrow from <italic toggle="yes">L</italic>4 to PS3 indicating that <italic toggle="yes">L</italic>4 was only used to define PS3 but is not a parent of <italic toggle="yes">A</italic>. (i) ADAG that denotes a propensity score PS4 for <italic toggle="yes">A</italic> as a function of <italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, and <italic toggle="yes">L</italic>4. (j) ADAG that depicts a propensity score PS5 for <italic toggle="yes">A</italic> specified as a function of <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>3, and <italic toggle="yes">L</italic>4, with a directed dashed arrow from <italic toggle="yes">L</italic>3 to PS5 indicating that <italic toggle="yes">L</italic>3 was used to define PS5 but is not an ancestor of A. The resulting data are unfaithful to the ADAG in that the independence of <italic toggle="yes">A</italic> and <italic toggle="yes">L</italic>3, conditional on PS5, is not evident using d-separation on the ADAG. <italic toggle="yes">L3</italic> should be avoided in estimating a propensity score because conditioning on PS5, a descendant of collider <italic toggle="yes">L3</italic>, induces a non-causal association between <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic>, which is evident from the ADAG.</p></caption><graphic xlink:href="nihms-2192435-f0001" position="float"/><graphic xlink:href="nihms-2192435-f0002" position="float"/><graphic xlink:href="nihms-2192435-f0003" position="float"/><graphic xlink:href="nihms-2192435-f0004" position="float"/><graphic xlink:href="nihms-2192435-f0005" position="float"/><graphic xlink:href="nihms-2192435-f0006" position="float"/><graphic xlink:href="nihms-2192435-f0007" position="float"/><graphic xlink:href="nihms-2192435-f0008" position="float"/><graphic xlink:href="nihms-2192435-f0009" position="float"/><graphic xlink:href="nihms-2192435-f0010" position="float"/></fig><table-wrap position="float" id="T1" orientation="landscape"><label>Table 1:</label><caption><p id="P18">Using augmented directed acyclic graphs (ADAGs) to assess variable selection,<sup><xref rid="TFN1" ref-type="table-fn">a</xref></sup> variable balance, and statistical efficiency for different propensity scores used in estimating the effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic>.</p></caption><table frame="box" rules="all"><colgroup span="1"><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/><col align="left" valign="middle" span="1"/></colgroup><thead><tr><th align="left" valign="top" rowspan="1" colspan="1">Propensity scores with different variable selections</th><th colspan="5" align="center" valign="top" rowspan="1">Variable balance before and after applying each propensity score (PS): whether <italic toggle="yes">A</italic> and each <italic toggle="yes">L</italic> are d-separated before and after applying the PS <break/><sup><xref rid="TFN2" ref-type="table-fn">b</xref></sup>Before applying the PS: <italic toggle="yes">A</italic> &#x022a5; <italic toggle="yes">L</italic> in DAG 1e<break/>After applying the PS: <italic toggle="yes">A</italic> &#x022a5; <italic toggle="yes">L</italic>|PS in ADAGs 1f to 1j</th><th align="center" valign="top" rowspan="1" colspan="1">Propensity score can be used to close open backdoors between <italic toggle="yes">A</italic> and <italic toggle="yes">Y</italic></th><th align="center" valign="top" rowspan="1" colspan="1">Can amplify uncontrolled confounding by including instrumental variable(s) in PS</th><th align="center" valign="top" rowspan="1" colspan="1">Statistical efficiency ranking (smallest asymptotic variance)<sup><xref rid="TFN3" ref-type="table-fn">c</xref></sup></th></tr></thead><tbody><tr><td align="left" valign="top" rowspan="1" colspan="1"/><td align="center" valign="top" rowspan="1" colspan="1"><italic toggle="yes">L</italic>0<break/>directly causes exposure only</td><td align="center" valign="top" rowspan="1" colspan="1"><italic toggle="yes">L</italic>1<break/>causes exposure and outcome</td><td align="center" valign="top" rowspan="1" colspan="1"><italic toggle="yes">L</italic>2<break/>causes exposure and outcome</td><td align="center" valign="top" rowspan="1" colspan="1"><italic toggle="yes">L</italic>3<break/>does not cause exposure or outcome</td><td align="center" valign="top" rowspan="1" colspan="1"><italic toggle="yes">L</italic>4<break/>directly causes outcome only</td><td align="left" valign="top" rowspan="1" colspan="1"/><td align="left" valign="top" rowspan="1" colspan="1"/><td align="left" valign="top" rowspan="1" colspan="1"/></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">PS1: <break/>P(<italic toggle="yes">A</italic>=1|<italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2)<break/>ADAG 1f</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS1: <break/>No<break/>After PS1:<break/>No<break/></td><td align="left" valign="top" rowspan="1" colspan="1">Before PS1: <break/>No<break/>After PS1:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS1: <break/>No<break/>After PS1:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS1: <break/>No<break/>After PS1:<break/>No</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS1: <break/>Yes<break/>After PS1:<break/>Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">No</td><td align="center" valign="top" rowspan="1" colspan="1">2 (second best)<break/>Reduces variance: <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2<break/>Magnifies variance: none</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">PS2: <break/>P(<italic toggle="yes">A</italic>=1|<italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2)<break/>ADAG 1g</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS2: <break/>No<break/>After PS2:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS2: <break/>No<break/>After PS2:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS2: <break/>No<break/>After PS2:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS2: <break/>No<break/>After PS2:<break/>No</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS2: <break/>Yes<break/>After PS2:<break/>Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">4<break/>Reduces variance: <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2<break/>Magnifies variance: <italic toggle="yes">L</italic>0</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">PS3: <break/>P(<italic toggle="yes">A</italic>=1|<italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4)<break/>ADAG 1h<break/></td><td align="left" valign="top" rowspan="1" colspan="1">Before PS3: <break/>No<break/>After PS3:<break/>No</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS3: <break/>No<break/>After PS3:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS3: <break/>No<break/>After PS3:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS3: <break/>No<break/>After PS3:<break/>No</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS3: <break/>Yes<break/>After PS3:<break/>Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">No</td><td align="center" valign="top" rowspan="1" colspan="1">1 (best)<break/>Reduces variance: <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4<break/>Magnifies variance: none</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">PS4: <break/>P(<italic toggle="yes">A</italic>=1|<italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4)<break/>ADAG 1i<break/></td><td align="left" valign="top" rowspan="1" colspan="1">Before PS4: <break/>No<break/>After PS4:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS4: <break/>No<break/>After PS4:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS4: <break/>No<break/>After PS4:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS4: <break/>No<break/>After PS4:<break/>No</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS4: <break/>Yes<break/>After PS4:<break/>Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">3<break/>Reduces variance: <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4<break/>Magnifies variance: <italic toggle="yes">L</italic>0</td></tr><tr><td align="left" valign="top" rowspan="1" colspan="1">PS5: <break/>P(<italic toggle="yes">A</italic>=1| <italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>3, <italic toggle="yes">L</italic>4)<break/>ADAG 1j</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS: <break/>No<break/>After PS:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS: <break/>No<break/>After PS:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS: <break/>No<break/>After PS:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS: <break/>No<break/>After PS:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Before PS: <break/>Yes<break/>After PS:<break/>Yes</td><td align="left" valign="top" rowspan="1" colspan="1">Opens a collider-biasing path between A and <italic toggle="yes">Y</italic> by conditioning on PS5, a descendent of collider <italic toggle="yes">L</italic>3</td><td align="center" valign="top" rowspan="1" colspan="1">Yes</td><td align="center" valign="top" rowspan="1" colspan="1">5 (worst)<break/>Reduces variance: <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2<break/>Magnifies variance: <italic toggle="yes">L</italic>0, <italic toggle="yes">L</italic>3</td></tr></tbody></table><table-wrap-foot><fn id="TFN1"><label>a</label><p id="P19">Variable selection using the backdoor criterion in DAG 1e: The minimum set of deconfounders is {<italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2}, which can close the open backdoors <italic toggle="yes">A</italic>&#x02190;<italic toggle="yes">L</italic>1&#x02192;<italic toggle="yes">Y</italic> and <italic toggle="yes">A</italic>&#x02190;<italic toggle="yes">L</italic>2&#x02192;<italic toggle="yes">Y</italic>.</p></fn><fn id="TFN2"><label>b</label><p id="P20">The symbol &#x022a5; is used to indicate &#x02018;independent of&#x02019;.</p></fn><fn id="TFN3"><label>c</label><p id="P21">An optimal or efficient valid adjustment set (e.g., <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, and <italic toggle="yes">L</italic>4 in DAG 1e) is the set of the parents of the outcome <italic toggle="yes">Y</italic>, excluding the consequences of exposure <italic toggle="yes">A</italic> that cause <italic toggle="yes">Y</italic>. This set contains the so-called &#x02018;precision variables&#x02019; (such as set <italic toggle="yes">L</italic>4) and yields the smallest asymptotic variance relative to the other variable adjustment sets in many standard parametric and non-parametric models [see, for example, Witte J, Henckel L, Maathuis MH, Didelez V. On efficient adjustment in causal graphs. <italic toggle="yes">J Machine Learning Res</italic> 2020;<bold>21</bold>:1&#x02013;45]. Including a parent (e.g., <italic toggle="yes">L</italic>0 in DAG 1e) of exposure <italic toggle="yes">A</italic> that is d-separated by <italic toggle="yes">A</italic> from <italic toggle="yes">Y</italic> worsens efficiency relative to the efficient valid set. Therefore, a PS model that includes the most parents of <italic toggle="yes">Y</italic> (without the parents of <italic toggle="yes">A</italic> only) would be expected to yield the smallest variance of the estimated effect of <italic toggle="yes">A</italic> on <italic toggle="yes">Y</italic> using a PS method. A PS-adjusted model, however, would yield a larger variance compared to using an outcome model that directly adjusts for <italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, and <italic toggle="yes">L</italic>4 (i.e., without a PS). This occurs because, when confounding equivalent sets exist for a DAG (e.g., PS3 and {<italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4} for ADAG 1h), adjusting for the variable set (<italic toggle="yes">L</italic>1, <italic toggle="yes">L</italic>2, <italic toggle="yes">L</italic>4) closest to the outcome <italic toggle="yes">Y</italic> tends to result in better statistical efficiency. Adjusting for variables closest (or those that are closer) to the exposure variable <italic toggle="yes">A</italic> tends to worsen statistical efficiency.</p></fn></table-wrap-foot></table-wrap></floats-group></article>