Transcript for 2026 Summer Training Series Session 5: Data Presentation and Visualization Presenter: Alexander F. Roehrkasse, Ph.D., Butler University National Data Archive on Child Abuse and Neglect (NDACAN) [VOICEOVER] National Data Archive on Child Abuse and Neglect. [ONSCREEN CONTENT SLIDE 1] Welcome to the 2026 NDACAN Summer training series! The session will begin at 12pm EST. This session is being recorded. Please submit questions to the Q&A box! [Alyssa Lindsey] All right, hello everyone. Welcome to the 2026 NDACAN Summer Training Series. My name is Alyssa Lindsey, I'm the Graduate Research Associate here at NDACAN. While folks are trickling in, and before we kick off the presentation, I just want to start with a couple of housekeeping items. We are recording this session so we can post the recording, the transcript, the slides, any related materials on our website a few weeks following each presentation. Those will be announced on the Child Maltreatment Research L, or the CMRL, listserv. If you want more information about that listserv and how to subscribe, I will put that in the chat, after my little spiel here. And please let me know if you have any questions by using the Q&A box. That is also where you can submit questions throughout the presentation, and we'll get to all of those questions as they come in at the end of the presentation. And that Q&A box can be found in the lower right-hand corner of the screen. Next slide. [ONSCREEN CONTENT SLIDE 2] NDACAN Summer Training series. National Data Archive on Child Abuse and Neglect Duke University, Cornell University, UC San Francisco, & Mathematica [Alyssa Lindsey] Great, so welcome again to the NDACAN Summer Training Series. NDACAN stands for the National Data Archive on Child Abuse and Neglect, and it's housed at Duke University, Cornell University, UC San Francisco, and Mathematica. Next slide. [ONSCREEN CONTENT SLIDE 3] Laying the groundwork foundational skills for using administrative data in child welfare research. Logo for the Children's Bureau features an image of overlapping blue and white silhouettes of children next to red and white stripes. To the right is the text "Children's Bureau: An Office of the Administration for Children & Families." Logo image which features a semicircle of icons representing people holding hands positioned on top of the acronym NDACAN, and next to the text National Data Archive Child Abuse and Neglect. [Alyssa Lindsey] NDACAN does two learning offerings every year. We have our monthly office hours series during the academic year, and in the summertime, we have the Summer Training Series. The theme of this year, 2026, Summer Training Series is "Laying the Groundwork, Foundational Skills for Using Administrative Data in Child Welfare Research". This series is designed for both the beginner, or folks that haven't used datasets before, or interacted with NDACAN datasets before, and also those who have, existing skill sets and are looking to strengthen their existing skills. So, all of the presentations throughout the training series are aiming to be a useful and thorough starting point for conducting research on child welfare using administrative data. Next slide. [ONSCREEN CONTENT SLIDE 4] NDACAN Summer Training series schedule. July 1st: Overview of NDACAN administrative datasets July 8th: Data cleaning and management July 15th: Linking NCANDS and AFCARS July 22nd: Handling missing data July 29th: Data presentation and visualization [Alyssa Lindsey] As I mentioned, this is the last, summer training series training, over data presentation and visualization. Our previous summer training series trainings included an overview of NDACAN administrative data sets, we had a session on data cleaning and management, and linking NCANS and AFCARs, and those first 3 sessions are up on our website. I'll put the link in the chat shortly. And then last week's training was over handling missing data. Next slide. [ONSCREEN CONTENT SLIDE 5] Session Agenda. Data presentation and visualization. Demonstration in R. Q & A. [Alyssa Lindsey] So again, today, our session is going to start with a presentation on data presentation and visualization, and then a demonstration in R, followed by Q&A. Again, feel free to put questions in the Q&A box as they come up for you, and we'll address all those questions in the order in which we receive them during the final Q&A session. Next slide. [Alyssa Lindsey] Now I'll pass it over to Alex to lead the rest of today's discussion. Thank you. [ONSCREEN CONTENT SLIDE 6] Data presentation and visualization. [Alex Roehrkasse] Thanks, Alyssa. My name's Alex Roehrkasse. I'm a research associate at NDACAN and an Assistant Professor of Sociology and Criminology at Butler University in Indianapolis. Thanks to everyone for being here today. I'm really excited for today's presentation, which, as Alyssa said, concludes our Summer Training Series. The presentation today is focused on data presentation, how to present data, or the results of an analysis of data, to an audience effectively. Mostly, we'll talk about audiences other than yourself, but as you'll see, we can also consider ourself an audience. How can we use data presentation to more effectively understand our own data? The analysis of our own data. We'll talk a lot today also about visualization. Visualization in particular as a method of data presentation. The slides today are going to focus mostly on first principles, best practices for data presentation and visualization. I'll talk through some canonical examples of effective and less than effective data presentation. The demonstration in R is going to be much more practical. I'll give you a little crash course in the most widely used packages for data visualization. And then talk through some very conventional ways that data are presented, particularly in academic journal articles. And discuss some alternative ways to present data and data analysis using data visualization. A reminder, as Alyssa said, that all of the materials for today's presentation are going to be posted on our website. And so, if you feel like things are moving a little too fast, don't worry, you can always go back, look back over a video recording of today's presentation. Also posted on our website will be all of the code that we talked through today. I invite you, I encourage you to explore that code, to use it, to adapt it, to whatever purposes you may need. Okay, let's get started. [ONSCREEN CONTENT SLIDE 7] Data presentation. Data presentation is foremost about communication. What is the message that you are trying to send? How will your audience receive your intended message? Data presentation is highly conventional and sometimes rule-bound, but it should be intentional and creative. With successful data presentation, tables and figures coordinate to communicate a coherent and convincing story. [Alex Roehrkasse] Well, first off, what is data presentation? What do we mean when we talk about data presentation? I think it's helpful to think about data presentation foremost as communication. Conceived as communication, data presentation invites a couple very important questions. First, What is the message that you are trying to send? When we think about communication, effective communication, we think about signals or messages that we're trying to send. Being clear with yourself about what this message is can make your data presentation more effective. Second, how will your audience receive your intended message? Those of us who have ever communicated with other human beings know that sometimes we have to communicate differently depending on our audience. Our audience may not understand our data in the same way that we do. It's important to use a little bit of imagination, even empathy, to try to anticipate how your intended audience will receive your intended message. An important thing to understand about data presentation is that it occurs in a context that is highly conventional. By this I mean, depending on your audience, there are often very strong conventions, or maybe even sometimes rules, about how data should be presented. Sometimes these conventions, these rules, are good. They're defensible. You should follow them and respect them. Sometimes, they're outdated, or were never really good ideas in the first place. And so one of the main messages I want to communicate in today's presentation is that when you think about data presentation. you should be intentional and creative. You should think for yourself about what is the most effective way to present your data or your data analysis. Lastly, when data presentation is successful, your tables and figures, which are the sort of fundamental categories of data presentation, should coordinate to communicate a coherent and convincing story. What this means is that any particular unit of data presentation, any particular table or figure, should itself be coherent and compelling. But you also want to think about how these different elements combine to tell a story. How each builds on the other, or responds to questions raised by the other. [ONSCREEN CONTENT SLIDE 8] Visualization as data presentation. Data visualization aids communication by visually encoding data. Visualization leverages pre-attentive attributes like size, color, or position. Compared to numerical information, pre-attentive attributes are processed more efficiently. Pre-attentive attributes have natural interpretations, e.g. “darker” is “greater.” Numerical information remains essential, but visualized quantities and other quantities of interest can be presented in appendix tables. [Alex Roehrkasse] Okay, what does it mean to think about visualization as data presentation? Data visualization fundamentally aids communication by visually encoding our data. What I mean by that is that through visualization, we give visual attributes to certain aspects of our data. Visualization is particularly powerful because it leverages what cognitive scientists call pre-attentive attributes. Things like size, color, position. Compared to numerical information, pre-attemptive attributes are processed more efficiently by our brains. Some of this is culturally specific, but much of it is not. These are hardwired into our brains as human beings. This is because many pre-attentive attributes have natural, intuitive interpretations. For example, most humans will see things that are darker and believe that they are greater relative to things that are lighter. The same is true of things that are larger. Larger is generally naturally, intuitively interpreted as greater. So too is higher. Our brains lend themselves to certain interpretations of pre-attentive attributes. We can leverage these attributes to make our communication more efficient, more clear. All that said, numerical information often remains really essential in a data presentation. We often do really need to see the numbers. But, we can often deliver these numbers in other ways. We can report visualized quantities and other quantities of interest in, say, a supplemental materials appendix, rather than in our main analysis. Thinking about coordinating between your main or primary data presentation and supplemental materials can be an important way to communicate effectively, but also transparently. [ONSCREEN CONTENT SLIDE 9] Data visualization as exploration. Image Description: A street map of London by John Snow showing the spatial distribution of cholera cases, where cases are plotted as points. The density of points is highest surrounding a water pump and decreases with distance from the pump. Source: Wikimedia Commons. [Alex Roehrkasse] Okay. As I said at the beginning of our presentation, we want to think about our audience. Most of today's presentation is going to be focused on other audiences, audiences other than yourself. But as a brief excursus, I want to suggest that data visualization isn't just about communicating the results of your analysis after you've finished it. Visualization is also really critical throughout the research process. And I'll give you two examples of that. First, data visualization as exploration, or part of the exploratory research process, and this usually comes early in your study. Let me give a really canonical example of data visualization as exploration. In the mid-19th century, there was a cholera outbreak in London. Scientists were trying to investigate the transmission mechanism of cholera. Competing hypotheses were that cholera was airborne or waterborne. A physician, John Snow, decided to approach the problem through visualization. What you see on the right is a map of the SoHo neighborhood, or part of the neighborhood, in London. Snow plotted as black Lines, stacked on top of one another, cases of cholera. He located those points where they occurred in that neighborhood. This yields a spatial visualization, of the spatial distribution of cholera cases. Intuitively, we see that there's a higher frequency of cholera cases at the center of this map. The frequency of cases decreases as we move further from the center of the map. This spatial pattern of cases is not consistent with an airborne transmission mechanism. In fact, what Snow found was the most cases were concentrated immediately adjacent to a frequently used water pump on Broad Street. Simply by visualizing the frequency of cases in physical space. We were able to develop a compelling test of alternative hypotheses. [ONSCREEN CONTENT SLIDE 10] Data visualization as diagnosis. Anscombe’s quartet. Source: Healy (2026). Image Description: Four Cartesian coordinate planes, each plotting eleven points and a fitted line. Despite the fitted lines being identical across the four plots, the distribution of points in each plot is very different. [Alex Roehrkasse] Data visualization can also be critical when we're evaluating our results, evaluating some hypothesis that we're starting to settle on. We might think of this as diagnosis, or sensitivity analysis, or robustness check. Let me give another very canonical example, what's called Anscombe's Quartet. On the right-hand side, you see visualized four different datasets. Each of these datasets has two variables. X values represented on the horizontal axis, Y values represented on the vertical axis. Each of these datasets is plotted on a Cartesian coordinate plane. And for each, we see a blue line, and this blue line represents the line of best fit, the line that captures the relationship between these two values, these two variables. What's hopefully immediately apparent Is that the blue lines are identical to one another. Each of the blue lines has the exact same slope, and the exact same intercept. This, despite the fact that, also quite obviously, the datasets are not the same. The relationships between X and Y in each of these datasets are not the same. If we were to run a regression model where we regressed Y on X, We would get the exact same model output for each of these four datasets. This would be misleading, because obviously the relationship between X and Y is very different in each of these datasets. What Anscombe's Quartet illustrates is that if we rely only on things like model statistics, say, regression coefficients, we may easily misunderstand the true relationship between different variables or other properties of our data. Data visualization is an important way of diagnosing our analysis, doing robustness checks on our findings. [ONSCREEN CONTENT SLIDE 11] Goals of visualization as data presentation. Capture and focus attention. Support intuitive, efficient, and accurate interpretation. Increase transparency and bolster credibility. [Alex Roehrkasse] Okay, moving then beyond ourselves as audience, let's return to this question of visualization as data presentation more generally, including for other audiences. At a most basic level, what are the goals of visualization as data presentation? Well, the visualization should capture and focus attention. I don't know about you, but I don't tend to find tables very visually striking. They don't draw in my attention. Figures, different data visualizations, often do, even when they're quite simple. Visualization can capture our attention, and then it can focus our attention on what's most important. Many data tables are often very large, they don't lend themselves to intuitive interpretation, and so visualization can help focus our attention relative to the tabular presentation of data. Visualization should support an intuitive, efficient, and accurate interpretation of our message. We should look at a visualization and say, yeah, that makes sense, I get it. We should be able to look at it and quickly understand what is the message that's trying to be sent. Lastly, an effective visualization should avoid misinterpretation. Finally, visualization can be especially important at communicating transparency and bolstering the credibility of our analysis. Most importantly, through the visualization of uncertainty. Most data and analyses of data involve some uncertainty. Visualization can be a particularly helpful way of visualizing that uncertainty, increasing the transparency and credibility of our findings. [ONSCREEN CONTENT SLIDE 12] Selected best practices for data visualization. Keep it simple. Label directly and comprehensively. Don’t trust the defaults. Avoid 3D. Always report uncertainty. Prioritize your message. Remember your audience. [Alex Roehrkasse] Okay, in light of these goals. What are some general best practices for data visualization? First, keep it simple. Sometimes a complex data visualization is beautiful, but is it an effective way of communicating? Very often, a complex visualization is better presented as multiple simple visualizations. Maybe if your visualization is very complex, you're trying to communicate three or four messages and it would serve you well to think about what are really the one or two most important messages you need to send with this visualization. Second, label things directly and comprehensively. Visualizations are most effective when the audience can see what different visual encodings mean in a quick and clear way. We'll talk about effective labeling in the demonstration a little later. Next, don't trust the defaults. As data visualization becomes more popular, more powerful, the software resources available to us in visualizing data become easier, more intuitive, but many of them include defaults, which may not serve our purposes effectively. Whenever you're visualizing data, consider carefully what your software is defaulting to, and whether or not there are ways to modify the defaults to more effectively create data visualizations. Almost always, you should avoid three-dimensional data visualization. Almost never do I see visualizations that need to be three-dimensional. Almost always, there's a more effective, creative way to visualize things more simply in two dimensions. Harkening back to our previous slide, we should always report uncertainty in our estimates, and visualization makes this easier and more intuitive. We should always prioritize our message over a beautiful visualization. And we should always consider how our audience expects to see certain things visualized. How we can leverage these expectations, navigate these expectations, to communicate clearly with our intended audience. [ONSCREEN CONTENT SLIDE 13] Misleading through ASPECT. Image Description: Three Cartesian coordinate planes, displaying identical scatter plots and fitted lines. Each plot has a different aspect ratio—one square, one stretched tall, one stretched wide—giving the false impression of diverging results. Source: Krause, Rennie, and Tarran (2024). [Alex Roehrkasse] Okay, having talked through some best practices, I want to illustrate a couple of the most common ways in which data visualization can go awry. The first is a seemingly trivial, but nevertheless common and consequential one. And this is what we sometimes call misleading through aspect. On the right-hand side, you see three different plots. Each of these plots visualizes the exact same data. The exact same points, the exact same fitted lines. The only difference between the three plots is that, as you can see, one is wider. One is taller, and the other is square. Simply by reshaping these plots, we give slightly different intuitive interpretations of what's really going on here. The left plot would seem to imply that the relationship here is somewhat gradual. The middle plot seems to imply that the relationship is dramatic. And the third is just a little more moderate. We don't want the conclusions or messages that our audience is receiving to be dependent on such trivial things as how we shape our visualization. For this reason, it's usually best to default to a square panel. And think carefully about whether and why any departures from this aspect are justifiable. Explaining why it's visualized in that way when appropriate. [ONSCREEN CONTENT SLIDE 14] Misleading through scale. Image Description: Two bar charts displaying the same information: ten random numbers ranging from eighty to one-hundred. The first bar chart has a vertical range of zero to one-hundred, while the second chart has a range of seventy to one-hundred. The latter plot gives the false impression of significant variation in the magnitude of plotted numbers. Source: Rougier, Droettboom, and Bourne (2014) [Alex Roehrkasse] Another common way in which visualization can mislead. This is misleading through scale. Again, we have the exact same dataset, visualized in two different ways. In both cases, we have 10 time periods arrayed horizontally, and visualized vertically the frequency of an event occurring in each time period. The frequencies range from 95 to 105. On the top, this is visualized with a vertical scale that ranges from 0 to 110. The bottom plot, plotted in red, shows only a partial range, ranging from 90 to 110. Neither is, strictly speaking, inaccurate, but because we fail to visualize the full range in the lower plot. It gives a mistaken impression of a very large degree of variation across time periods. Simply put, the size of the bars seems to change a lot. The visual encoding of the frequency, and the variation in frequencies, is misleading. The upper plot gives a more intuitive sense that, yes, the frequency varies, but as a proportion of the whole quantity, not very much. We want to think carefully about how questions of scale can mislead, and this is important when thinking about defaults. Many statistical software packages, packages for visualizing data, will default to a partial range. We need to think carefully about manually specifying a full range when this is appropriate to avoid misleading through scale. [ONSCREEN CONTENT SLIDE 15] Presenting comparisons effectively. When plotting variables with different scales, avoid dual-axis charts and instead use multiple panels. When reporting regression coefficients for variables with different scales, use standardization or log-transformation. When visualizing ratio variables, scale axes logarithmically. When analyzing heterogeneity, plot predicted values or marginal effects rather than reporting interaction terms. [Alex Roehrkasse] Often, when we're presenting data, we're trying to make comparisons. Compare one group to another, compare one model to another. Visualization can be a particularly helpful way of making comparisons. But there's other things that can frequently go awry whenever we're trying to present a comparison. Whenever we're plotting variables with different scales, it's important to avoid dual-axis charts, instead using multiple panels. I'll illustrate this in the demonstration shortly. If we're reporting regression coefficients for different variables that have different scales. Insofar as we intend to compare the magnitude of these effects, these coefficients, we need to make sure that we're standardizing or transforming our coefficients so that they're meaningfully comparable. If we're visualizing a ratio variable, it's important to scale our axes logarithmically. This is to allow us to compare directly ratios larger and smaller than 1. For example, the difference between one half and 1 should be visually equivalent to the difference between 1 and 2. Lastly, whenever we're analyzing heterogeneity, we should, instead of reporting interaction terms, which are not intuitive at all, very frequently very difficult to interpret. Instead, we might choose to present predicted values or marginal effects that give much more intuitive understandings of how one relationship depends on the values of another variable. [ONSCREEN CONTENT SLIDE 16] Using color carefully. A screenshot of the ColorBrewer interface. The interface allows for the user to select the number of data classes, the nature of the data (sequential, diverging, or qualitative), and the desired color scheme. A map illustrated the color scheme and the interface clarifies the specific colors used and whether the scheme is colorblind safe, print friendly, or photocopy safe. Source: ColorBrewer (colorbrewer2.org). [Alex Roehrkasse] Color is one of the most exciting and compelling aspects of data visualization, but it's also an area where things can go a little bit awry. Many of the default color schemes in visualization packages can mislead or may not be appropriate to the message you're trying to send. The Colorbrewer is a very helpful resource for developing color schemes that can more effectively communicate your intended message. Colorbrewer was developed by a number of American cartographers who were originally trying to think about colorized maps, but Colorbrewer is helpful for all kinds of visualization using color. Colorbrewer can help you develop a color palette depending on the nature of your data, whether it's sequential, or diverging, or categorical. It can help you select color palettes that are accessible for people with colorblindness, make sure that your visualization will work if it's photocopied or printed. Colorbrewer is a fundamental, a foundational resource for data visualization today. [ONSCREEN CONTENT SLIDE 17] Ensuring Accessibility. Outline colorized objects. Ensure contrast (4.5:1 for text, 3:1 for objects). Label directly. Consider multiple encoding. Use alt-text. [Alex Roehrkasse] Following directly from this, visualization presents a lot of opportunities to make our data presentations more intuitive, more efficient, more clear. But it also has trade-offs, particularly in the area of accessibility. Making sure that visually impaired audiences can understand our message, can receive our message, is imperative. There are certain strategies for data visualization that can help ensure that our visualization is accessible. Some of these include outlining colorized objects, making sure that our visualization includes sufficient contrast. Again, labeling things as much as possible, and as directly as possible. Encoding things in multiple ways, by which I mean, if some distinction depends on color alone, it might be more accessible to encode a distinction not only using color, but also shape. Or color and line type. I'll demonstrate this in the presentation in the demonstration shortly. Lastly, using alt text is an important way for people to be able to, read or hear a description of your data presentation as visualization. [ONSCREEN CONTENT SLIDE 18] Related considerations. Pre-registration. Supplemental materials . Replication packages. Open access and pre-prints. [Alex Roehrkasse] Lastly, there are a few other things, strictly speaking, beyond the scope of data visualization, but which go to effective data presentation. They're beyond the scope of today's presentation, but I want to highlight them briefly. Pre-registration, where before conducting your analysis, you report publicly what your analysis plan is, can be a helpful way of increasing the transparency and credibility of your message. Using supplemental materials, appendices, can be a helpful way of reporting robustness checks, or data underlying your visualizations. Publishing replication packages can be a way for other people to replicate your visualizations, investigate the data underlying them, or extend your approach to data presentation to other cases. Lastly, publishing your research open access is an obvious way to increase the impact of the messages you send. [ONSCREEN CONTENT SLIDE 19] Demonstration in R. [Alex Roehrkasse] With that, we'll start to move to a demonstration in R. As I said earlier, this demonstration will be much more focused on the practical how-tos of data presentation, a crash course in what we call ggplot, and also a discussion of some conventional ways of tabular data presentation, some visual alternatives to those conventions. [ONSCREEN CONTENT SLIDE 20] Additional Resources. Fundamentals of Data Visualization, Clouse O. Wilke. Data Visualization (2e), Kieran Healy. R Graphics Cookbook (2e), Winston Chang. [Alex Roehrkasse] Before we move to R, though, as I have in other presentations, a few additional resources for extending our learning today. Each of these is an open source resource, freely available online. All three of these I highly recommend for extending an exploration of data visualization, particularly with applications in R. Okay, with that, we'll now move to our demonstration in R. We'll be working in Rstudio today. If you've been with us for any of our other presentations, or you are yourself an Rstudio user, hopefully this environment will be familiar. We'll be working from an R script today, specific to this presentation. [ONSCREEN CONTENT] NOTES This program file demonstrates strategies discussed in session 5 of the 2026 NDACAN Summer Training Series "Data presentation and visualization." For questions, contact the presenter Alex Roehrkasse (aroehrkasse@butler.edu). Note that because of the process used to anonymize data, all unique observations include partially fabricated data that prevent the identification of respondents. As a result, all descriptive and model-based results are fabricated. Results from this and all NDACAN presentations are for training purposes only and should never be understood or cited as analysis of NDACAN data. [Alex Roehrkasse] A few brief notes, as with other weeks. We'll be working today with data that's been fabricated so as to anonymize the identities of people captured in NDACAN data. For this reason, the data we'll be discussing, the analysis of it, should not be understood or cited as analysis of archived data. The data we're working with today have realistic properties, but are not real data. [ONSCREEN CONTENT] TABLE OF CONTENTS 0. SETUP 1. INTRO TO GGPLOT2 2. DESCRIPTIVE STATISTICS 3. MODEL STATISTICS [Alex Roehrkasse] We'll mostly skip over our setup for today, but see previous presentations about how to set up your working environment. And then we'll focus today's presentation on, first, a crash course in ggplot2, the package widely used, most widely used in modern data visualization in R. And then we'll talk through, different ways to present descriptive statistics and model statistics. So, very often, table 1 in an academic journal article will be a kind of table of summary statistics describing your data. We'll talk about how to create such tables, but also how to create alternative visualizations. Our main results, particularly in a statistical analysis, are often model statistics. Again, we'll talk through briefly some ways to present these as tables, but then also some alternative approaches To visualize model statistics. [ONSCREEN CONTENT] > # Let's clear the environment. > rm(list=ls()) > # Pacman installs packages if necessary, otherwise loading them. > if (!requireNamespace("pacman", quietly = TRUE)){ + install.packages("pacman") + } > pacman::p_load(data.table, tidyverse, + mice, + knitr, + ggstance, gtsummary, modelsummary, scales) > # Let's define some filepaths (note the organization of project and data folders). > project <- 'C:/Users/aroehrkasse/Box/Presentations/-NDACAN/2026_summer_series/' > data <- 'C:/Users/aroehrkasse/Box/NDACAN/2026_summer_series/' [Alex Roehrkasse] Okay. We'll set up the environment by first clearing any objects in our environment, and installing or loading packages of interest. Note today that we're, using the ggplot is part of the tidyverse package. We're using the mice package to work with some imputed data. And then a number of other packages for visualization. These are all packages I frequently use to do tabular or visual presentation of data. [ONSCREEN CONTENT] > # And set one as the working directory. > setwd(project) > > # Always set a seed to allow for reproduction of random processes. > set.seed(1013) > # Let's read in our cleaned, linked data: children 0-3 entering foster care in 2023 (AFCARS) linked to maltreatment histories (NCANDS) (see session 3). > dlink <- read_rds(paste0(data,'linked_data.rds')) > # Let's also load our complete-case and multiply imputed model estimates of substantiation/indication (see session 4). > load("saved_models.RData") [Alex Roehrkasse] We'll otherwise set up our environment and read in two kinds of data today. The first dataset we'll read in is a cleaned, linked dataset that we created in session 3. These are children 0 to 3 entering foster care in 2003, which we then linked to maltreatment histories from ndans. In session 4, we then imputed missing data in this linked dataset, and so we will read in some saved models based on those imputed data. Models of the likelihood of substantiation or indication, particularly in an ncans sample. A reminder that today's presentation is just based on some limited data from New England and a small number of variables, just so that we don't get bogged down in computation today. [ONSCREEN CONTENT] > # The modern grammar of data visualization in R is implemented in the ggplot2 package, part of the Tidyverse. > # To create a ggplot, you first apply the function to a dataset and define aesthetics. Notice, though, that we still don't have any data visualized. > p1 <- ggplot(dlink, + aes(x = nrep, + y = nsub)) > p1 [Alex Roehrkasse] Okay, first, a brief introduction to ggplot2, the sort of grammar of graphics in data visualization today in R. Mostly, this is implemented through the ggplot2 package, which is part of the tidyverse. What is a ggplot, and how do we create it? First, we take the ggplot function. And we apply it to a dataset. In this case, our linked dataset. We also, in order to create a ggplot, need to define some aesthetics. We do this using the aes argument. Here, we'll define two aesthetics. The x-axis, or the horizontal axis, will represent the number of reports, number of maltreatment reports that each child has experienced before reaching age 3. The vertical axis will be the number of substantiated reports that a child has experienced. We'll go ahead and run this function, and in so doing, create a gzplot object, p1. If we then visualize or plot this object, we see over here on the right, a cartesian coordinate plane. With our variables of interest, but obviously no data. What's going on? Well, this is because a ggplot object on its own is not a visualization of data. It's almost just like an environment in which we can then start to visualize data. We actually begin to visualize data when we add geoms to our ggplot. We do this simply with the addition sign and start creating any number of different geoms. [ONSCREEN CONTENT] > # This is because we then have to add geoms. > p1 + + geom_point() Warning message: Removed 189 rows containing missing values or values outside the scale range (`geom_point()`). [ONSCREEN CONTENT R CODE IMAGE 1] Graph with x-axis nrep (number of maltreatment reports for a child) and y-axis nsub (substantiated reports). Points show integer data, but each point is a large number of observations, and therefore the visualization does not illustrate the data distribution well. [Alex Roehrkasse] So we take our ggplot object, and we add to it a point geom. Okay, now we're starting to see data actually visualized in our cartesian coordinate plane. This is not a very good visualization. It's honestly not very helpful at illustrating anything. Because our data are integer data, it's very difficult to get a sense of the distribution of the data, because each of these points actually represents a very large number of observations of children having that exact combination of values. [ONSCREEN CONTENT] > # We can then modify these geoms by adding arguments to them. Note how each approach (imperfectly) leverages pre-attentive attributes. > p1 + + geom_point(alpha = .05) # set transparency Warning message: Removed 189 rows containing missing values or values outside the scale range (`geom_point()`). [Alex Roehrkasse] So, what we can do is modify this geom. By adding arguments. Which tweak or modify the aspects of the genome. One way we might do this is by adjusting the transparency of the points. So that if it's a small number of points, it appears fairly transparent, but if a large number of points overlay one another, it starts to seem darker. [ONSCREEN CONTENT R CODE IMAGE 2] A plot with x-axis nrep (number of maltreatment reports for a child) and y-axis nsub (substantiated reports). Points are faded to indicate lower counts. [Alex Roehrkasse] Okay, now we're getting somewhere. Because we intuitively understand that darker is greater, we see that in this area, we're dealing with a larger number of observations. And a smaller number of observations out here, where the points are much fainter. There are other ways to approach this, though. [ONSCREEN CONTENT] > p1 + + geom_point(position = position_jitter(width = .2, + height = .2)) # place noisily Warning message: Removed 189 rows containing missing values or values outside the scale range (`geom_point()`). [ONSCREEN CONTENT R CODE IMAGE 3] A jitter style plot with x-axis nrep (number of maltreatment reports for a child) and y-axis nsub (substantiated reports). Points are arranged in clumps to visually indicate counts/frequency. [Alex Roehrkasse] Another is to tell R to kind of shake up the location of these points on the cartesian coordinate plane, or to jitter the position. Now, instead of the points perfectly overlaying one another, we get sort of an area of points where we can intuitively see that, okay, there's a lot of observations over here, and not so many observations over here. Different ways of visually encoding frequency. [ONSCREEN CONTENT] > # Furthermore, we can overlay multiple geoms. > p1 + + geom_point(position = position_jitter(width = .2, + height = .2)) + + # A line illustrating identity between the two axes + geom_abline(slope = 1, + linewidth = 1, + linetype = 'dashed') + # accessible encoding + # A locally estimated line of best fit + geom_smooth() `geom_smooth()` using method = 'gam' and formula = 'y ~ s(x, bs = "cs")' Warning messages: 1: Removed 189 rows containing non-finite outside the scale range (`stat_smooth()`). 2: Removed 189 rows containing missing values or values outside the scale range (`geom_point()`). [ONSCREEN CONTENT R CODE IMAGE 4] A ggplot2 graphic with three geoms 1) a jitter style plot with x-axis nrep (number of maltreatment reports for a child) and y-axis nsub (substantiated reports), 2) a dashed line with slope of one overlaid to indicate whether the number of reports is equal to the number of substantiated reports, and 3) a locally estimated line of best fit [Alex Roehrkasse] The next step in creating a ggplot is to overlay multiple geoms. So let's say we stick with this jittering approach, we can then add another geom, for example, a straight line with a slope of 1, indicating whether or not the number of reports is equal to the number of substantiation reports. Furthermore, we can create another geom, geomsmooth, that is a locally estimated line of best fit. Okay, we see now that we have 3 geoms, the points. The dashed line, and the smooth line, all overlaying our ggplot. But it's pretty difficult to know what's going on here, and so this is where labeling and cleaning up or thinking more creatively about our aesthetics can really start to enhance our message. [ONSCREEN CONTENT] > # But to really know what's going on, we need to label things clearly and completely. > p1 + + geom_point(position = position_jitter(width = .2, + height = .2)) + + # Make color an aesthetic so it can be included in a legend + geom_abline(aes(intercept = 0, + slope = 1, + color = '100% substantiation'), + linewidth = 1, + linetype = 'dashed') + + geom_smooth(aes(color = 'Locally fitted line')) + + # Set colors manually + scale_color_manual(values = c('100% substantiation' = 'red', + 'Locally fitted line' = 'blue')) + + # Label axes and suppress legend title + labs(x = 'Number of maltreatment reports', + y = 'Number of\nsubstantiated reports', + color = NULL) + + # Choose theme + theme_classic() + + # Design legend + theme(legend.position = c(0.01, 0.99), + legend.justification = c("left", "top"), + legend.background = element_rect(color = 'black')) `geom_smooth()` using method = 'gam' and formula = 'y ~ s(x, bs = "cs")' Warning messages: 1: Removed 189 rows containing non-finite outside the scale range (`stat_smooth()`). 2: Removed 189 rows containing missing values or values outside the scale range (`geom_point()`). [Alex Roehrkasse] Let's define as an aesthetic of the diagonal line it's intercept, slope, and color. This will allow us to assign it a legend. Let's also assign, as an aesthetic of the smooth line, its color. We'll then set our colors manually, so that one line will be red, and one will be blue. We'll label all of our axes, choose a more compelling theme. And then design a legend that will communicate what is being plotted. [ONSCREEN CONTENT R CODE IMAGE 5] A ggplot2 graphic with a legend (“100% substantiation”, “Locally fitted line”), labeled axes (x-axis is “Number of maltreatment reports”, y-axis is “Number of substantiated reports”), and colored lines. [Alex Roehrkasse] If we create this plot, now we're really starting to see a clearer message emerge. While this figure is far from perfect, there's a lot I would take issue with, what we can see is that it starts to communicate a couple of messages more clearly. First, most children have a small number of reports and substantiated reports. But a smaller number of children did have a larger number of reports. Third, the proportion of reports that were substantiated decreased with the number of reports. We can see this as the blue line of best fit. Starts to diverge from the red dashed line, indicating 100% substantiation. The distance between these lines grows as the number of maltreatment reports grows. A somewhat subtle message communicated intuitively by a simple visualization. Okay, let's keep going now and talk through a couple different, very conventional ways of presenting data, particularly in an academic journal article, and explore some alternative ways of presenting that data through visualization. [ONSCREEN CONTENT] > # DESCRIPTIVE STATISTICS > # Very conventional in an academic journal article is a table of descriptive statistics. Take, for example, our linked dataset. Some helpful canned functions can generate nice-looking descriptive tables, e.g. the gtsummary package. > dlink |> + tbl_summary(include = c(nsub, ageatlatrem)) [ONSCREEN CONTENT R CODE IMAGE 6] A basic descriptive table generated by gtsummary that shows the frequency and proportional frequency of each value of each variable (nsub is number of substantiated report, ageatlatrem is age at last removal). [Alex Roehrkasse] Very conventional is to have a table of descriptive statistics. Let's look at our linked dataset. And we might use a canned command. In this case, a command from the gtsummary package that's designed to very simply summarize the descriptive properties of some data. So we take our linked data object. And we apply the summary command to it. And gtsummary very happily outputs, you know, a decent-looking table that shows the frequency and proportional frequency of each value of each variable. It even shows the number of missing values. [ONSCREEN CONTENT] > # Sometimes, though, manual calculation is more flexible. > dlink_sum <- dlink |> + select(nsub, ageatlatrem) |> + summarize(across(everything(), + list( + Mean = ~mean(., na.rm = T), + Median = ~median(., na.rm = T), + SD = ~sd(., na.rm = TRUE), + P10 = ~quantile(., 0.1, na.rm = T), + P90 = ~quantile(., 0.9, na.rm = T) + ))) |> + pivot_longer(cols = everything(), + names_to = c('var', '.value'), + names_sep = '_') |> + mutate(var = case_when(var == 'ageatlatrem' ~ 'Age', + var == 'nsub' ~ 'Substantiated investigations')) > dlink_sum # A tibble: 2 × 6 var Mean Median SD P10 P90 1 Substantiated investigations 1.36 1 0.877 1 3 2 Age 0.827 0 1.05 0 3 [Alex Roehrkasse] Sometimes, though, it can be helpful to manually calculate some of these quantities of interest. So let's instead select two of these variables of interest, the number of substantiated reports and the age at last removal. We'll summarize across these two variables. And in doing so, generate a few different summary statistics. A couple measures of central tendency, the mean and median, a measure of the variance, the standard deviation. And then, a couple other measures of the distribution of the data, the 10th and 90th percentile quantities. We'll do a little reorganization and relabeling to create this d-link summary object. We see this object pop up in our environment. And if we inspect the element. We see that really all it is is a very simple table, where we have a clearly labeled variable, its mean, median, standard deviation, and percentiles. [ONSCREEN CONTENT] > # We can save this table as a CSV or XLSX file, or output it as a table in LaTeX using the knitr package. > kable(dlink_sum, format = "latex", booktabs = TRUE) \begin{tabular}{lrrrrr} \toprule var & Mean & Median & SD & P10 & P90\\ \midrule Substantiated investigations & 1.3648000 & 1 & 0.8772849 & 1 & 3\\ Age & 0.8267014 & 0 & 1.0492420 & 0 & 3\\ \bottomrule \end{tabular} [Alex Roehrkasse] We can then export this table as a csv or an excel file. We can even output it for latex users into a table that can be copied and pasted into a latex document, say, on overleaf. Okay, this would be a very conventional way of presenting descriptive statistics. Let's consider, however, how visualization might combine some of the merits of each of these types of tables, intuitively presenting both the full distribution of the variables that we see in the categorical table, but also presenting some information about the central tendencies. [ONSCREEN CONTENT] > # Consider, however, how visualization can combine some of the merits of each of these types of tables, intuitively presenting both the full distribution of the variables and their central tendencies. > dlink |> + select(nsub, ageatlatrem) |> + pivot_longer(cols = everything(), + names_to = 'var', values_to = 'num') |> + mutate(var = case_when(var == 'ageatlatrem' ~ 'Age', + var == 'nsub' ~ 'Substantiated investigations')) |> + ggplot() + + # Plot a histogram with width of 1 + geom_histogram(aes(x = num), + binwidth = 1) + + # Use our summarized data from above to plot central tendencies + geom_point(data = dlink_sum |> + pivot_longer(cols = c('Mean', 'Median'), + names_to = 'stat'), + aes(x = value, + # Multiply encode the central tendencies + color = stat, + shape = stat), + y = 0, + size = 3) + + # Design scales + scale_x_continuous(breaks = 0:10) + + scale_y_continuous(labels = label_comma()) + + # Choose an accessible color palette + scale_color_brewer(palette = 'Set2') + + # Create separate panels for each variable + facet_wrap(~var, scales = 'free_x') + + theme_bw() + + guides(color = guide_legend(reverse = T), + shape = guide_legend(reverse = T)) + + labs(x = NULL, y = 'Number of\nobservations', + color = NULL, shape = NULL) + + theme(legend.position = 'bottom') Warning message: Removed 189 rows containing non-finite outside the scale range (`stat_bin()`). [Alex Roehrkasse] We can take our linked data object, reorganize it a little bit, and then plot it. Notice that we haven't defined any aesthetics in our ggplot. We'll instead do that within each specific geom. So, in our ggplot, we'll first plot a histogram. This histogram will have a width 1, we'll then use our summary data object that we created a little bit ago, reorganize it a little bit, and also plot it on the horizontal axis. We'll multiply encode these central tendencies so that both color and shape distinguish the mean from the median. We'll design our horizontal and vertical axes, choose an accessible color palette using the color brewer function, and we'll facet wrap our plot to create separate panels for each variable, rather than plotting them on dual axes. [ONSCREEN CONTENT R CODE IMAGE 7] The number of observations for the two variables “substantiated reports” and the “age at last removal” are shown in separate histograms with central tendencies median and mean indicated by shape and color. [Alex Roehrkasse] When we run this command we see a pretty simple, pretty intuitive, summary, of our data. The histogram communicates the full distribution of our data, intuitively and clearly. The points show the central tendencies. They're clearly distinguished in terms of both color and shape, and we can see which is which in a clear legend. Note that the vertical axis of both panels is identical, so that we can compare the frequency of values directly. But we've allowed the horizontal axis to vary because the range is different for the two variables. There are obviously trade-offs in presenting your descriptive statistics in this way. I mean only to highlight that visualization can be a creative, intentional way of presenting data unbound of conventions for the presentation of descriptive data. [ONSCREEN CONTENT] # 3. MODEL STATISTICS > # Recall from session 3 that we estimated a basic logistic regression model using complete-case analysis and multiple imputation. Conventionally, we would report these results as a table. Again, there are some helpful canned functions for table generation (cf. the stargazer package). > modelsummary(list(m_cc2, pool(m_mice))) [Alex Roehrkasse] What about model statistics? Let's say we run a regression model. Very common is to report a table of regression statistics. Recall that from a previous session, we estimated two models in which we predict the likelihood of substantiation or indication of a maltreatment report as a function of a child's prior frequency of reports and, one other, predictor yes, whether they experienced financial distress. We estimated two models, one using a complete case analysis, where we dropped observations with missing values, and another where we imputed missing values multiple times. [ONSCREEN CONTENT R CODE IMAGE 8] A basic descriptive table generated by the modelsummary command. It summarizes two logistic regression models from Session 3 in which we predict the likelihood of substantiation or indication of a maltreatment report as a function of a child's prior frequency of reports and whether they experienced financial distress. The table contains coefficients, standard errors, a number of different statistics reporting, and goodness of fit. [Alex Roehrkasse] We can use a canned command, like model summary, to, again, output a, you know, a decent-looking table, where we have the two models here, the coefficients, standard errors, a number of different statistics reporting, goodness of fit. This table, though, is not publication-ready. Happily, many of these canned commands make it easy to customize your tables, adding, removing, or labeling information to make our message clearer. How might we do this? [ONSCREEN CONTENT] > # These packages make it easy to customize your tables to add/remove/label information to make your message clearer. > modelsummary(list('Complete case' = m_cc2, + 'MICE' = pool(m_mice)), + statistic = "conf.int", + conf_level = 0.95, + stars = T, + coef_omit = "Intercept", + coef_rename = c("chpriorYes" = "Prior report", + "fcmoneyYes" = "Financial distress"), + gof_omit = ".*") [ONSCREEN CONTENT R CODE IMAGE 9] A regression-results table comparing Complete case and MICE analyses for "Prior report" and "Financial distress". Confidence intervals are shown in brackets. Footnotes are denoted by the plus sign, and by groups of stars (*, **, ***). [Alex Roehrkasse] Well, we might name the different models, so that instead of model 1, model 2, we have, okay, very clearly, this is a complete case analysis, this is an analysis based on missing, multiply-imputed data. Instead of reporting standard errors, which are often difficult to interpret or make sense of, we can report confidence intervals, which much more clearly help us understand whether our, say, 95% confidence interval spans zero. We can include stars, essentially footnotes that help us understand the statistical significance of our results. We can, say, suppress coefficients that are not of interest, rename variables, or suppress goodness-of-fit tests that don't actually enhance the message we're trying to send. This is a much clearer, more effective way of presenting regression model statistics in a tabular format. [ONSCREEN CONTENT] > # While this table is clear, it's not altogether intuitive. Consider how visualization can strenghten our evidence-based message. > # First, we extract, clean, and combine quantities of interest from our models, creating a plot_data data frame. > cc_summary <- summary(m_cc2) |> + coef() |> + as.data.frame() |> + rownames_to_column('term') |> + mutate(model = 'Complete case') |> + rename(est = Estimate) |> + select(term, est, model) > cc_ci <- confint(m_cc2) |> + as.data.frame() |> + rownames_to_column('term') |> + rename(lower = `2.5 %`, upper = `97.5 %`) Waiting for profiling to be done... > cc_combined <- left_join(cc_summary, cc_ci, by = 'term') > mice_combined <- pool(m_mice) |> + summary(conf.int = TRUE) |> + as.data.frame() |> + mutate(model = 'MICE') |> + rename(est = estimate, lower = `2.5 %`, upper = `97.5 %`) |> + select(term, est, lower, upper, model) > plot_data <- bind_rows(mice_combined, cc_combined) |> + mutate(term = factor(term, + levels = c('(Intercept)', + 'chpriorYes', + 'fcmoneyYes'), + labels = c('Intercept', + 'Prior report', + 'Financial distress')), + est = exp(est), + lower = exp(lower), + upper = exp(upper)) |> + filter(term != 'Intercept') [Alex Roehrkasse] While I think this table is clear, particularly to audiences who are experienced in reading regression tables, it's not altogether intuitive. So let's finally consider how visualization can strengthen our evidence-based message. First, we'll do a little bit of organizing our data, where we will extract from our models some quantities of interest, clean and combine them? I won't talk through this code, but I invite you to explore how it works when we post these codes on our website. [ONSCREEN CONTENT] > # We can then plot the quantities in the plot_data data frame. > plot_data |> + # Define aesthetics shared across geoms + ggplot(aes(x = est, + y = fct_rev(term), + color = model, + # Multiply encode for accessibility + shape = model, + group = model)) + + # Highlight the null hypothesis as a visual anchor + geom_vline(xintercept = 1, linetype = 'dashed') + + # Plot point estimates as points + geom_point(position = position_dodgev(height = -.5), + size = 2) + + # Plot confidence intervals as error bars + geom_errorbarh(aes(xmin = lower, + xmax = upper), + height = .25, + position = position_dodgev(height = -.5)) + + # Log-transform axis of ratio estimand for comparison above/below 1 + scale_x_continuous(trans = 'log', + breaks = seq(.5,2.25,.25), + limits = c(.5, 2.25)) + + # Choose a colorblind-safe color palette + scale_color_brewer(palette = 'Dark2') + + # Clearly label all plot features + labs(x = 'Odds ratio', y = 'Predictor', + color = 'Missing data\nstrategy', + shape = 'Missing data\nstrategy') + + ggtitle('Substantiation/indication of\nCPS reports') + + theme_bw() + + theme(axis.text.x = element_text(angle = 45, + hjust = 1, + vjust = 1.1), + legend.position = 'bottom') Warning messages: 1: position_dodgev requires non-overlapping y intervals 2: Using the `size` aesthetic with geom_path was deprecated in ggplot2 3.4.0. ℹ Please use the `linewidth` aesthetic instead. This warning is displayed once per session. Call lifecycle::last_lifecycle_warnings() to see where this warning was generated. [Alex Roehrkasse] Once we've cleaned, organized, combined our model estimates, we can plot them. We'll go ahead and take this data object for creating our plots. We'll create a ggplot in which we define aesthetics shared across all geoms. I'll go ahead and first create this plot, and then talk through the code that creates it. [ONSCREEN CONTENT R CODE IMAGE 10] Forest plot titled “Substantiation/indication of CPS reports” comparing odds ratios for two predictors under two missing-data strategies: complete case (teal circles and confidence intervals) and MICE (orange triangles and confidence intervals). A dashed vertical line marks an odds ratio of 1.0. For “Prior report,” the complete-case estimate is below 1 (about 0.85), while the MICE estimate is above 1 (about 1.6). For “Financial distress,” the complete-case estimate is above 1 (about 1.55), while the MICE estimate is near 1.0. Horizontal error bars show confidence intervals. [Alex Roehrkasse] Notice that we have some shared aesthetics across the plot. The x-axis is our estimates. The y-axis is our regression terms. Color and shape both encode our different models. The multiple encoding allows our audience not to rely solely on color, but also on shape. We highlight the null hypothesis, namely that our odds ratio is 1, by drawing a dashed line at 1. This way, we know that any confidence interval that crosses over this line fails to reject the null hypothesis. We then plot our point estimates as points, and our confidence intervals as error bars. Note that we have log-transformed our x-axis, this is so that say, a coefficient of one half is visually equivalent to a coefficient of 2. We've chosen a colorblind safe color palette, and we've clearly labeled all features of our plot. The plot itself has a title, the coefficients are clearly labeled, our estimate is clearly labeled. We have a labeled legend, with the two groups clearly labeled. This is an example of how we might report regression statistics, or model statistics, visually. We get a much more intuitive sense of our results than we do using a regression table. Of course, I would strongly encourage you to report a table of regression estimates in a supplemental materials appendix. But as you try to develop a coherent and compelling story in your main data presentation, visualization can help you send your message more effectively to your intended audience. [ONSCREEN CONTENT SLIDE 21] Questions? User support: NDACANsupport@cornell.edu Garrett baker: garrett.baker@duke.edu Alyssa Lindsey: Alyssa.lindsey@ucsf.edu [Alex Roehrkasse] Okay, this concludes our demonstration, and so I'll return to our slide deck. Reminding you that we are always here for questions. I hope that we can use our remaining 8 minutes to talk through any immediate questions about the presentation, but any questions that don't arise today, arise this afternoon, tomorrow, next month, next year, please don't hesitate to email our general support line, or me personally. We're always happy to support users of our data with questions specific to our data, or more general questions about foundational skills for working with archive administrative data. Thanks, everyone. I look forward to your questions. [Alyssa lindsey] Thank you, Alex. I put our emails in the chat for everyone. I'm also going to put in the chat our Frequently Asked Questions page on our website, and that can also be a helpful resource as you, review the data yourself. So just a reminder, if you have questions, please put them in the q and a box, it's in the lower right-hand corner of your zoom screen. And we'll wait a few minutes to take any questions that come in. If not, we might wrap up. All right, I'm not seeing any questions come in, and I see folks are hopping off, so maybe this is a good place to end. Thank you all so much for being here today. Again, this is our last session of the Summer Training Series, and all these resources will be uploaded in the next few weeks. Alex, if you have any additional things to add. [Alex roehrkasse] No, I just want to thank everyone for your time and attention, for joining us this summer, for contributing your questions. At the risk of repeating myself, I'll just say that we eagerly welcome outreach from data users with questions. So please do visit our site to explore the code, other aspects of the presentations that we'll be archiving shortly. But as you continue working with archived data, please don't hesitate to reach out with any questions, requests for support. [Alyssa Lindsey] Great. Thanks, all! [Alex Roehrkasse] Thanks, everyone, take care. [VOICEOVER] The National Data Archive on Child Abuse and Neglect is a joint project of Duke University, Cornell University, University of California, San Francisco, and Mathematica. Funding for NDACAN is provided by the Children's Bureau, an office of the Administration for Children and Families. [Music]