Working Paper Series Congressional Budget Office Washington, D.C. CBO's Simulation Model of New Drug Development Christopher P. Adams Congressional Budget Office christopher.adams@cbo.gov Working Paper 2021-09 August 2021 To enhance the transparency of the work of the Congressional Budget Office and to encourage external review of that work, CBO's working paper series includes papers that provide technical descriptions of official CBO analyses as well as papers that represent independent research by CBO analysts. Papers in that series are available at http://go.usa.gov/ULE. Thanks to David Austin, Anna Anderson-Cook, Margaret Blume-Kohout, Joseph DiMasi, Ru Ding, Pierre Dubois, Michael Falkenheim, Craig Garthwaite, Sebastien Gay, Ryan Greenfield, Tamara Hayford, Manuel Hermosilla, Evan Herrnstadt, Benedic Ippolito, Jeffrey Kling, Ellen Werble, Chapin White, and James Williamson for helpful suggestions. Thanks also to Katherine Feinerman for reviewing the code, Gabe Waggoner for editing, and Erik O'Donoghue and Jorge Salazar for help in developing the figures. www.cbo.gov/publication/57010 Abstract This paper presents the Congressional Budget Office's simulation model for analyzing leg- islative proposals that may substantially affect new drug development. The model uses estimates of changes in expected future profits or development costs to estimate the per- cent change in the number of drug candidates entering the various stages of human clinical trials. Given changes in decisions to enter at each stage, the model estimates when and by how much the number of new drugs entering the market will change. To illustrate the implications of the model, the paper considers a legislative change that lowers expected returns for the top-earning drugs. A 15 percent to 25 percent reduction in expected returns for drugs in the top quintile of expected returns is associated with a 0.5 percent average annual reduction in the number of new drugs entering the market in the first decade under the policy, increasing to an 8 percent annual average reduction in the third decade. The analysis takes the estimated impact of the policy on expected returns as given. In CBO's assessment, those estimates are in the middle of a wide distribution of potential effects. The effects could be smaller if expenditures in late-phase human trials are larger, for example. Alternatively, the effects could be larger if the cost of capital is larger. Keywords: health care, prescription drugs, new drug development JEL Classification: 111, 118 The model is based on a stylized representation of the pharmaceutical decisionmaking process. A firm is projected to continue development of a drug if expected returns exceed expected costs. The model's parameter values are derived from both estimation and cali- bration procedures. The model uses revenue estimates calculated using nonpublic Medicare Part D data. Those data include information on the rebates paid by the manufacturing firms, allowing those rebates to be netted out of revenue. Using data from 2010 to 2018, CBO estimates how drug revenue varies with time on market. CBO uses results from Di- Masi, Grabowski, and Hansen (2016) to estimate development costs. The agency uses a Roy model to combine information on revenue and cost. That model accounts for selection and correlation in the observed revenue and cost data (Heckman and Honoré 1990). Although an input into the Roy model is observed revenue, the output is expected returns. The as- sumption of "revealed preference" is used to elicit the firm's expectations about the value the firm will receive from bringing the drug to market. Those expectations account for var- ious costs associated with producing, selling, and distributing the drug-even though those costs are not observed in the data. That said, CBO does have access to rebate information in the Medicare Part D data and in the estimation procedure nets out manufacturer-paid rebates. CBO calibrated entry probabilities and other parameters on the basis of results presented in Blume-Kohout and Sood (2013), DiMasi (2013), and Khmelnitskaya (2020). The analysis shows that any relationship between a policy change and the number of new drugs entering the market grows over time. The change would be small for the first few years because key decisions for drugs entering in those years would have been made before the policy change. However, the size of that change would increase substantially as decisions in earlier phases of development affect later phases. The estimates of Dubois and colleagues (2015) and Acemoglu and Linn (2004) can be thought of as the effect averaged over time. Blume-Kohout and Sood (2013) and Dranove, Garthwaite, and Hermosilla (2020) esti- mate how increases in market size affect drug development over time. Both papers describe the impact of introducing Medicare Part D, called the Medicare Modernization Act. Both research groups use pipeline data to show how increases in market size affected the num- bers and types of drugs entering each phase of development. Both reports show that the impact of the changes increases over time. Dranove, Garthwaite, and Hermosilla (2020) use a longer panel and information on the novelty of the drug to show that the initial effect is on increasing development of the least novel drugs. That paper shows that the policy change took many years to affect the entry of the most novel drugs into various phases of development. Two mechanisms determine the observed change in entry into a particular phase of development. First is an immediate change in whether a potential drug will enter one of the three phases of development. Second are changes to the candidates available to enter a phase of development given changes to earlier phases of development. The Blume-Kohout and Sood (2013) estimate of a 27 percent immediate increase in phase I trials stems from the first mechanism, and the long-term effect of a 50 percent increase stems from both mechanisms combined. Similarly, the authors find that the initial impact on phase III trials is small but becomes much larger as decisions from earlier phases show up in changes in the number of phase III trials. Dranove, Garthwaite, and Hermosilla (2020) show an initial impact of non-novel drugs entering preclinical and clinical development; by definition, those are the drugs available to enter development. Drugs that are more novel take longer to go through the process, taking longer to become available for entry into the different phases. 2 Background on Drug Development IND NDA/BLA Application Submitted Design and Clinical Trials Exploratory (5 to 7 Years) Preclinical Studies >'@r> @dr@)> -@ 726 > 6 Process e Phase | Development @& (1 to 3 Years) Preclinical, Regulatory Toxicology Phase II Review by FDA Studies ee (1 to 2 Years) (2 to 4 Years) wou ies ee @ ee Phase Ill "agen (2 to 3 Years) Figure 1: Evaluation and research chart of the drug development process. BLA = Biologic License Application; FDA = Food and Drug Administration; IND = Investigational New Drug; NDA = New Drug Application. Drugs go through a systematic and regulated process illustrated in Figure 1. During the initial preclinical development period, preliminary scientific research is done to determine what type of drug may work on the disease. That period also includes studies of the drug working in animals, potentially including those genetically modified, to help assess how the disease affects humans. After determining an appropriate drug candidate, a company may enter its drug in human clinical trials. Human clinical trials generally begin with an Investigational New Drug application to the Food and Drug Administration (FDA) or equivalent international agency. Human trials follow a regulated and standardized three-phase process. In general, a drug candidate 3.1 Hurdle Model Before deciding to enter the phase (phase k € {1,2,3}), the firm observes a signal of the drug candidate i's "type," denoted 4;,. That signal gives the firm expectations over the drug candidate's likelihood of success, returns once on the market, and costs of development. The firm will enter the phase of development if net expected returns are positive conditional on the observed signal. The firm will enter phase & with drug candidate ¢ if and only if the following inequality holds: E(yin Ri, - Cii,|9ix) > 0 (1) where yz; € {0,1} indicates whether the drug candidate will successfully complete the phase, Rj, are the expected returns associated with successfully completing the phase, and CZ, are the expected costs associated with entering the phase. Equation (1) states that the firm will enter the phase if and only if the expected returns exceed the expected costs. The asterisk means that those values are not necessarily observed in the data set. The identification and estimation issues are discussed below. 3.2. Dynamic Model Each decision is linked in that the following relationship holds: Riqg_1) = E (max{0, yin Ri, - Ci} 10ie-1)) (2) Conditional on the observed signal (6;(,_1)), the decisionmaker knows both the expected costs and the expected return of the current phase, where the second value includes the expected costs of the next phase. That said, the decision in the next phase is based on a new draw of the signal. A drug expected to get a low return in phase IT may end up having a high return in phase III. In the notation above, the signal at the beginning of phase k -1 (9:(h-1)) is independent of the signal at the beginning of the next phase (9;,). Although that parameterization is less realistic, it substantially simplifies modeling and estimation. Note also that expected net returns from the next phase may be negative. Expected costs of the phase may exceed expected returns from entering the phase. However, the firm knows that it does not have to enter the phase if net expected returns are negative. Entering phase & - 1 has an option value. The firm can choose not to take the drug into phase k if information available at that time suggests the drug will be unprofitable. 3.3. Simulation Model The simulation includes a static part and a dynamic part. The static part simulates the decision to enter a particular phase of development. That part takes the number of available drug candidates as given and determines which will enter the next stage of development. The dynamic part simulates the interaction of the static decisions. That part accounts for how decisions made in earlier development stages affect the number of available candidates in later stages. The dynamic part also models the time candidates take to become available for the next stage. The static decision considers three values of interest for drug candidate i entering phase k: whether the drug will successfully complete the phase, y;,; expected return from completing the phase, Rj; and expected costs of entering the phase, Cj,. The simulation proceeds by drawing those three values for many pseudo-drug candidates for each phase of development. In the model, y;; is independent of the other two values. That independence implies that the parameter p;, = Pr(yiz = 1) can be directly estimated from drug development pipeline data such as those presented in DiMasi, Grabowski, and Hansen (2016). In the model, the other two values are distributed by bivariate distribution. {log( in)» log (Ci) } ~ F (Urcks Orcky Prek) (3) where [rx is a vector of the mean of the expected log return and the mean expected log costs, Oreck is the equivalent vector for standard deviation of log expected returns and log expected costs, and p;-x% captures the correlation across expected returns and costs. The bivariate distribution (F) is given by a Gaussian copula function where the expected cost marginal is a log-normal and the expected return marginal is a log-gamma distribution. The gamma distribution can capture the skewness of the data. As discussed below, the parameterization allows the distribution to be estimated with the data available. The full simulation model is dynamic. At each period (a year), a set of candidates is available to enter each of the three phases of development. For each candidate, the values described above are drawn and used to determine whether the candidate enters the development phase. That decision, and the probability that the candidate will successfully complete the phase, determines whether the candidate is available to enter the next phase. Drug candidate i's time in phase k& is represented by t;,. As with other values, that time is assumed to be distributed log-normal and determined by parameters 4, and oy. The amount of time and the expenditure (F;,,) in the development phase are modeled using a bivariate log-normal distribution, where {log(t;,), log(ix)} ~ N (Hier, Ztex), Mtek = {Lek Lek}, and 2 _ otk PtekO tk? ek ek PtekOtkO ek o 2 (4) is the variance-covariance matrix for observed time in development and expenditure in development. In general, those values are observed and taken from survey data presented in DiMasi, Grabowski, and Hansen (2016). Those two bivariate distributions are related through the interaction between expected costs of development and observed expenditure in development. 4 Identification The model of the firm's decision problem has several key inputs: estimates of success probabilities, expected development costs, expected time in development, financing costs, and expected returns. Because CBO doesn't observe all those values, the agency calibrates some parameters of the model and estimates others with restrictive parameterizations. This section discusses identifying and estimating the joint distribution of expected returns and costs for each phase of development. 4.1 Roy Model CBO uses information from two sources to estimate parameters of the model. For expected returns, the agency uses Medicare Part D data that describe what is paid to manufacturers of individual drugs over their lifetime. Those data include confidential information on rebates paid by manufacturers. For costs, CBO uses results from a survey of pharmaceutical manufacturers reported in DiMasi, Grabowski, and Hansen (2016). Two concerns arise from using that information to estimate the parameters. First, ex- pected returns and expected costs are not observed as a pair for each drug candidate. Rather, CBO observes the marginal distribution of returns for one set of drug candidates and the marginal distribution of costs for another set. Second, the data set suffers from selection bias. CBO doesn't observe expected returns for drug candidates under consider- ation to enter development. Instead, the agency observes returns for drugs that actually entered and later completed development and then entered the market. Both problems are solved by modeling how observed distributions are related to distributions of interest through the decision problem presented above. In particular, this working paper uses a parametric Roy model to identify the joint distribution of expected returns and expected costs (Heckman and Honoré 1990).° Consider a simple version of the problem, with one data set in which expected returns for drugs take on two values, high and low. In a second data set, expected costs for developing drugs take on two values, high and low. Those two data sets contain only returns and costs for drugs that enter development. (For simplicity, assume that returns and costs are observed for all drugs that enter development.) For the firm, eventual returns and costs are known before deciding to take the drug candidate into development, which occurs only if expected returns exceed expected costs. If the cost data set describes the proportion of drugs with high costs, what can be learned about returns for those drugs? They must have high expected returns. If not, the drugs wouldn't be observed in the data set. If the cost data describe drugs with low costs, what can be learned about expected returns for those drugs? Not much, because drugs with low costs would be developed, and appear in the data, for both high and low expected 3In the standard parametric Roy model, the econometrician observes the marginal distributions of out- comes in the two sectors, relative prices and the market share of the two employment sectors. returns. Because their costs are low, those drugs are observed in the data set regardless of expected returns. If the returns data set describes the proportion of drugs with low returns, what can be learned about costs? Those drugs must have low expected costs. Again, otherwise the drugs wouldn't be observed in the data set. However, observing the proportion of drugs with high returns does not allow inferences about expected drug costs. Those drugs enter development with both high and low expected costs. The two data sets describe the proportion of drugs with high costs and high returns as well as the proportion of drugs with low costs and low returns. Given that and with knowledge of the proportion of drugs that enter development, the two data sets can be used to infer the proportion of drugs with high costs and low returns and the proportion of drugs with low costs and high returns. In that simple problem, the returns and cost data sets supply enough information to determine all the joint probabilities even though neither data set includes both returns and costs for any particular drug. In the more general model, there exists an expected return Rj and an expected cost C7 and a decision problem in which the firm enters if and only if Rf - C7 > 0. Let R; and C; represent observed values for returns and costs. By assumption, those values are observed only if the profitability condition holds. Moreover, those observed values are the expected values known to the firm at the time of the choice. __ Sf RE ifRt-cr>o Ry -{ - otherwise (5) c -{& if Ri Ch >0 (= - otherwise In addition, CBO observes the probability of entering Pr( Rj - C7 > 0) = 7. Heckman and Honoré (1990) show that this model is not nonparametrically identified unless the "price" changes.* Here, the "price" is 1 and does not change. That paper also shows that requiring the distribution to be a bivariate normal allows the parameters to be identified. Given that result, CBO uses a similar distribution for {R*,C*}.° Following the approach of Heckman and Honoré (1990), CBO assumes that expected returns and expected costs observed by firms are equal to observed returns and costs after the filtering process is accounted for.® The model used below is more complicated. In particular, the joint distribution has four marginals: success rate, returns, costs, and time in development. Given that, several "In the original example, the price is the relative wages in the two sectors. ®In the actual Roy model, the econometrician observes C; when RY - CO <0. ®Heckman and Vytlacil (2007) discuss identification issues with a more general model that allows dif- ferences between values observed by the firm and values observed in the data. In the policy simulation presented below, CBO considered a robustness check wherein the firm observes an additional signal about the relative profitability of the drug candidate. Though not presented, those results show that the impact of the policy is attenuated. restrictions are placed on the model. The success probabilities for each phase are assumed to be independent of the other factors. Moreover, phase success rates observed by the decisionmaker are assumed to be equal to observed success rates presented in DiMasi, Grabowski, and Hansen (2016). In addition, the observed costs are the capitalized expenditures. Those values are deter- mined by expenditures during the phase, time in the phase, and a discount rate. Therefore, CBO needs to estimate the joint distribution of expenditures and time in development. DiMasi, Grabowski, and Hansen (2016) present the marginal distribution for expenditures and time, not the joint distribution. As a result, CBO calibrates the correlation across expenditures and time, comparing results with those presented in DiMasi, Grabowski, and Hansen (2016). See discussion below. 4.2 Capitalized Expenditures DiMasi, Grabowski, and Hansen (2016) describe the actual expenditure in each phase of development by drug. But those amounts do not account for actual costs of development because they don't account for the opportunity cost of the money. Money invested in human clinical trials for a particular drug candidate could have been invested in some other project. Capitalized expenditures are calculated in two steps. First, expenditures are assumed to be uniformly spread out over the period. That spread is approximated by rounding time in phase to the nearest year and assuming that equal fractions are spent at the beginning of each year. Those amounts are discounted to the end of the phase and summed. Second, capitalized expenditure in the phase is discounted to the expected time of entry on the market. All dollar amounts are discounted to the point where the drug enters the market. That discounting is done to correctly compare various expenses that occur during the drug's development process and the returns that the drug receives while on market. tin-l Cx = ( S- (*) (1+ ay) (1+ B) Ter) (6) fog \ bik where Ej, is the observed expenditure of drug i in phase k, @ is the discount rate (time cost of money), t;, denotes time in phase k, and T;,, denotes time to market from the beginning of phase k. Again, CBO doesn't observe the joint distribution of expenditures and time in de- velopment. Therefore, the agency uses a log bivariate normal distribution and calibrates the correlation parameter to other statistics presented in DiMasi, Grabowski, and Hansen (2016). In particular, those authors summarize the actual expenditure, time, and capital- ized expenditures in the phase. With the discount rate that study used, the correlation parameter is the one that most closely matches the summary statistics presented in the 10 paper. 5 Data The model uses two main sources of data. Returns data come from Medicare Part D expenditures on brand-name drugs. Cost information comes from DiMasi, Grabowski, and Hansen (2016). 5.1 Estimates of Returns To estimate the model, CBO needs estimates of the distribution of returns for drugs entering the U.S. market. To do that, CBO uses Centers for Medicare & Medicaid Services data on Medicare Part D expenditures from 2010 to 2018. That data set includes confidential information on manufacturer-paid rebates, allowing them to be netted out. Those drugs are a convenience sample. They also are relevant to policy proposals of interest. Unfortunately, that data set does not describe returns from the rest of the U.S. market or the global market. Therefore, CBO calculated Medicare Part D's share of global revenue and then used that percentage to estimate annual global revenues for each drug. That estimating approach implicitly projects that the proportion of drugs with large returns (and small returns) is the same in Part D and globally. That approach does not, however, require an assumption that a given drug's share of revenue is the same in Medicare Part D and globally. Rather, the shape of the distribution is what matters. CBO regresses returns on a polynomial of age, the number of years since the launch of the drug in the United States. rig = a(t - to) + a(t - tos)? + 03(t - toi)? + vit (7) In that equation, to; is the launch year of drug ¢, rz is the observed returns of drug 7 in year t, and vj are unobserved characteristics of drug returns. Equation (7) is a cubic without an intercept term. The actual regression also includes a time trend for the year in which returns are observed. The cubic parameterization is used to capture drug life-cycle returns. Returns tend to start low as prices and market share are low and then increase as both go up over time. Returns then start to fall as older drugs face greater competition from the entry of brand- name and generic drugs (DiMasi, Grabowski, and Vernon 2004; Bhattacharya and Vogt 2003). To estimate equation (7), CBO combines the data set for returns with information on the launch date of the drug. (Returns data are calculated at the ingredient level.) To estimate a distribution of returns, CBO uses quantile regression. Coefficient estimates and implied revenue by year and percentile are available in supplemental data posted with this working paper. To compare costs and returns on the same basis, each drug's returns are 11 log-scale of $ related papers further discuss the survey design. Development expenditures used here do not include expenditures associated with marketing or postmarket studies. Phase | Mean | Median SD I 25.3 17.3 29.6 II 58.6 44,8 50.8 Til 255.4 200.0 153.3 Table 1: Expenditure distributions in millions of dollars for each phase, from DiMasi, Grabowski, and Hansen (2016). SD = standard deviation. Table 1 shows a skewed distribution of expenditures. In the supplement to their 2016 paper, DiMasi, Grabowski, and Hansen present fitted log-normal distributions for expen- ditures and time in phase. As mentioned, to estimate capitalized expenditures of drug development, CBO needs to know the joint distribution of expenditures and time in development. That information is not presented in DiMasi, Grabowski, and Hansen (2016). Therefore, CBO calibrates the correlation parameter by using various measures of distributions of expenditure and time in development presented in the paper. The calibration exercise gives a correlation coefficient of -0.5. As with other parameters, sensitivity around that value is discussed and tested below. Concern exists that survey results of expenditures presented in DiMasi, Grabowski, and Hansen (2016) are not representative of drugs in development. In Adams and Brantner (2006) and (2010), the authors use publicly available data and compare estimates with those in DiMasi, Hansen, and Grabowski (2003), which used a similar analytic method to that in DiMasi, Grabowski, and Hansen (2016). Using development times and transitions from Pharmaprojects and R&D expenditure from Securities and Exchange Commission filings, Adams and Brantner present results broadly similar to those in DiMasi, Hansen, and Grabowski (2003). CBO doesn't use the discount rate presented in DiMasi, Grabowski, and Hansen (2016); rather, the agency uses the WACC for pharmaceutical and biotech industries suggested by Damodaran (2020).® 5.3 Other Observed Distributions As inputs, the model also uses estimates for success rates of drugs through clinical trials and amount of time spent in each phase. Here CBO uses reported results in DiMasi, Grabowski, and Hansen (2016). Those numbers are from a sample of more than 1,400 drug candidates initially tested in humans between 1995 and 2007. Similar results are presented in Wong, ®Data used to calculate the WACC were from the January 2020 update. As with other measures, the sensitivity of the results to that parameter value is discussed more below. 13 Siah, and Lo (2019) for more than 15,000 candidates in development between 2000 and 2015. In addition, CBO sets the probability of entering each phase of development to .1, .2, and .9 for phases I, II, and III, respectively. Those numbers are chosen mostly because they lead to feasible results. CBO doesn't have access to information on which drugs are available to enter a particular phase but are not doing so. Because of concern about those values, this working paper presents information on how much the estimates vary if those values change. A survey by DiMasi (2013) asked whether the failure of the drug in development was due to safety, efficacy, or commercial reasons. He found that commercial reasons were less prominent in later phases than in earlier phases, a pattern consistent with the assumption. In recent work, Khmelnitskaya (2020) found that 8.4 percent of all attrition is strategic. 6 Estimation The object of interest is the joint distribution of expected returns and expected costs ({ Ri, Ci}). However, because CBO cannot observe those values, the agency observes val- ues that have gone through the decisionmaking "filter" described above, {R;, C;}. To esti- mate model parameters, CBO solves a simulated generalized method of moments (GMM) problem. GMM is a generalization of a standard least squares problem. A moment is the term used to describe the average of some set of values raised to a power. The first moment is the average of values to the power 1, the second moment is the average of values to the power 2, and so on. Consider the problem of finding the average in a sample. In expectation, the difference between the sample value and the average is zero. Alternatively, the average in the sample is where the first moment is zero. Similarly, the sample variance is where the second moment is zero (when the average is zero). Different moments characterize different aspects of the distribution. The problem is that with multiple moments, potential exists to have multiple estimates of the same parameter. The question thus arises about which estimate to use or how to average across various estimates. Hansen (1982) presents properties of a particular method of weighting across moments, the GMM. Because CBO can't observe the underlying joint distribution of expected costs and expected returns, the agency must simulate it. Firms decide whether to enter simulated drugs into development. Simulated expected costs and expected returns for drugs that go into development are then compared with those that are observed. 6.1 Simulated GMM The estimator works through a choice of the five parameters-{purk, Ork; ck, Tcky ANA Prck }- that characterize the distribution of expected returns and costs ({log(R%,), log(Cj,)}). Given that choice, the estimator simulates many expected return and expected cost pairs. 14 Phase | Variables Bb m a S p x | Log Shift I Revenue 1.14 |} 9.59 | 4.18 | 10.80 | 0.93 | 0.10 | -1.00 Cost 2.68 | 3.35 | 0.64 | 0.94 -5.36 II Revenue 2.80 | 8.26 | 6.70 | 10.49 | 0.82 | 0.20 | --1.00 Cost. 4.05 | 4.43] 0.62 | 0.75 -15.54 III Revenue 5.53 | 4.70} 9.05 | 3.21 | 0.98 | 0.90 | --1.00 Cost 3.69 | 5.36] 7.09] 0.93 -9.97 Table 2: Parameter estimates and moments from the simulated generalized method of moments estimator. Estimated distributions of expected costs and expected returns are represented by a mean of yz and standard deviation of o. Expected returns are distributed log-gamma and expected costs are distributed log-normal. The correlation across expected costs and expected returns is represented by p. The mean and standard deviation of ob- served costs and returns are represented by m and s, respectively. The probability of entering the phase of development is 7. The "Log Shift" column refers to the fact that observed distribution is shifted up to use logs. Using logs simplifies the estimation, but the values must be shifted to avoid zeros or negative numbers. Phase | Variables LB 0 p Mean | SD | Mean} SD | Mean} SD I Revenue 1.70 | 0.99 | 7.92 | 2.45 | 0.83 | 0.19 Cost 2.48 | 0.35] 0.77 | 0.06 II Revenue 2.96 | 0.87 | 10.71 | 5.05 |} 0.83 | 0.07 Cost 3.91 | 0.21 | 0.68 | 0.06 Il Revenue 6.94 | 7.58 | 13.42 | 21.37] 0.98 | 0.01 Cost 3.52 | 1.39] 8.29] 4.33 Table 3: Variation in parameter values as a result of uncertainty. Estimated expected cost and expected returns distributions are characterized by the following parameters: mean, p; standard deviation, 0; and correlation, p. Mean and standard deviation (SD) of parameter estimates are based on sampling uncertainty associated with expected costs and expected returns estimates and on variation in assumed parameter values, such as probability of entering the trial. 16 presented below are based on estimates presented in Table 2, although variation around entry effects is presented in section 8. As discussed further below, uncertainty over the estimates comes from several sources. One concern is that uncertainty is associated with sampling variation in both the returns estimates from the Centers for Medicare & Medicaid Services data and the cost and du- ration estimates from DiMasi, Grabowski, and Hansen (2016). A second concern is that several parameter values were set using limited information. 7 Policy Impact: An Illustration A range of policies could be analyzed. In particular, any policies that significantly affect expected returns or expected costs for new drugs may lead to changes in the number of drugs that get to market. To illustrate how the model works, CBO considers a policy that significantly reduces expected returns of drugs in the top 20 percent of expected returns. That is, drugs expecting to land in the top quintile would generate expected returns 15 percent to 25 percent less than without the policy. That is a representative policy that affects expected returns similarly to the one proposed in H.R. 3. The policy then required the Secretary of Health and Human Services to negotiate drug prices and prioritize drugs to areas where the impact would be greatest. The bill also capped the price at which parties could negotiate. The price could not exceed 120 percent of an international price index. CBO estimated that the policy would decrease future global revenue for new drugs by 19 percent (CBO 2019a, 2019b). In the main analysis, this working paper considers how a policy that reduces expected returns for the drug affects new drug development. The appendix considers an additional impact of a policy that reduces the cash available to invest in new drug development. 7.1 Impact on Phase III Decisions CBO assumes that the policy's impact is increasing over the distribution of expected re- turns. At the top end of the distribution is a 25 percent reduction in expected returns. That reduction falls to 15 percent for a drug expected to be at the 80th percentile and then to zero reduction in expected returns below the 80th percentile. Figure 3 illustrates the impact of the policy. Expected returns are not affected for any drug below the 80th percentile. The figure shows that the policy has a small impact on the number of new drugs entering phase III. The policy affects only a few drugs on the margin between entering and not entering phase III. Only a few simulated drugs are near the line, and the change in expected returns is not large enough to have many drugs cross that line. CBO estimates that the policy is associated with a 0.6 percent decrease in the number of drugs entering phase II]-that is the immediate impact. Below, this paper discusses the longer-term impact of decisions made earlier in the R&D process. 17 * Baseline Policy Expected Return rn. B100M S108 Expected Cost 7.2 Impact on Phase IT Decisions The impact of the policy on phase II decisions is filtered through the phase III decision process. The first step is to determine the effect of the policy on the net expected returns for phase III (expected returns less expected costs) from successfully completing phase II. Comparing the distribution of net expected returns with and without the policy, CBO finds a 25 percent reduction in expected returns for drugs with net expected returns above the 90th percentile, a 20 percent decrease for drugs with net expected returns between the 75th and 90th percentiles, and a 5 percent decrease for drugs with net expected returns below the 75th percentile.® Figure 4 presents the impact of the policy on the distribution of phase II expected returns and costs. The policy causes expected returns to fall. Sometimes that fall is enough to move the drug from above to below the 45-degree line. The figure shows that the policy affects more marginal drug candidates. Here it leads to an immediate 3 percent decrease in the number of drugs going into phase II. 7.3 Impact on Phase I Decisions For phase I decisions, the impact of the policy is filtered through the decision problems for both phase IT and phase IIT. When the net expected returns with the policy are compared with net expected returns at the baseline, the policy is associated with a 25 percent decrease in expected returns. That effect leads to an immediate 4 percent reduction in the number of drugs entering phase I. 7.4 Impact of the Policy Over Time Figure 5 shows that the impact of the policy initially grows before leveling out. The policy is estimated to reduce the number of drugs coming to market. A 0.5 percent decrease occurs in the first decade under the policy, a 5 percent decrease in the second decade, and an 8 percent decrease in the third decade. The changes over time in the simulated impact of the policy are partly due to two effects. First, the model allows firms to remove drugs already in development if the reduction in expected returns falls below the costs associated with remaining in the development phase. Second, in the model, the policy affects the number of drugs entering each phase of development. The analysis above shows that the policy has little effect on drugs entering phase III, with larger effects on the number of drugs entering phases II and I. Those decisions further back in the development process accumulate over time, with fewer drugs moving from phase I to phase II and then fewer still moving into phase III. °CBO does not explicitly estimate those findings; rather, they are assumed based on observed changes from the simulations. 19 * Baseline Policy c = a = a e oD a = o a a = Lu BOOM $200M Expected Cost 46 Baseline g = HA AM HK HH ae 9g KK HI I 3 IH DCN Kae IE Hy = *o = s i= o @ Te . ao oS * = e E a Palicy '5 ', a "Sane « a bz os Sete g08,8 Sasa ao Se = = oo om 0 5 10 15 20 25 30 40 Years After Policy Implementation Figure 5: The policy would not affect the number of drugs entering the market in the short run but is expected to have long-run implications. The policy is implemented in year zero, but the full difference is not reached until after year 20. To illustrate, the number of new drugs is initially set at the average for 2015-2019. 8 Uncertainty The results presented here are uncertain. Uncertainty exists around both the values of inputs used in the simulation model and the impact of the illustrative policy. Using the illustrative policy, the section shows how uncertainty over the model's input values and inherent uncertainty in the simulation affect predictions of the policy's impact on the number of new drugs. For example, if the WACC is higher, the policy has a larger effect on the number of new drugs. Conversely, if expenditures in phase III are higher, the policy has a smaller effect on the number of new drugs. The distributions shown in Figure 6 account for uncertainty over the exact value of parameters set by CBO and uncertainty over the exact value of parameters estimated outside the simulation model. For input values set by CBO, a uniform distribution of 21 values is used, ranging from a "small" decrease to a "small" increase in the parameter value. The exact size varies, but for probabilities it is generally 10 percentage points. For estimates coming from distributions presented in DiMasi, Grabowski, and Hansen (2016), a bootstrap procedure is used in which the sample size is the one equal to the survey sample size for each phase. For estimates of returns, the quantile regressions are bootstrapped. Phase Ill Phase Il 0.00 0.04 0.02 0.03 0.04 0.05 Reduction in Probability of Entering Phase Figure 6: The policy's impact decreases as the drug moves from phase I to phase III. The uncertainty over the policy's impact is much higher for phase I and phase II than for phase III. One reason is that earlier phases use estimates from later phases. That is, estimates for the impact on phase I and II entry incorporate the uncertainty for the impact on phase III entry. In this paper, CBO estimates that the policy would lead to an immediate decrease in the percentage of drugs entering each phase of development. Figure 6 presents distributions for those estimates. For phase III, the estimated impact of the policy is very small, and the 19 A set of paramcters estimated in DiMasi, Grabowski, and Hansen (2016) are treated as set outside the estimation procedure used here. Those are the probabilities of completing each phase and the duration from the end of the phase to market. 22 figure shows that variation around the estimate is relatively tight. The policy's effect on entry into phase II and phase I is larger, as is variation around the estimate. That greater variation occurs at least partly because more estimated parameters are associated with those values. The policy's impact on phase II entry is determined by how the policy affects the distribution of expected returns and how that change is filtered through the phase III decision problem. x = x we o xs xR as = x 4 x * x aot a * Pa x . 4 a Mowe * xx x x * x * * = a * = 4325 ° * x _ 'o * es e AK RK . '4 "of a REX: OK CK j Ko Mie ee xe ox = oO x x x . + ic a a & bo io x a ae em # exe eK ae y= a Kk Ke OR OX) OM Ke Ke ER ekee ox oe £ x ee CO x eee =e Re REPRE KX Mtxt id 2 ee x 2 axe eee Shits fe \MaRe ane ERE = Eee Sep yg ERE ERR een alee nyn ae ees S xX eX xh EERO SS otonneeseegoser topes =| a ry = Kewx XR # a + & Eaue Bac ee eee taent ee OX EX _ = xe eK Ff FF eee eee FRE o xe =~ xte x x* HK HM FF Ke x o eK m [ex HX x® Rese a me ee a x xt * * * * x = tXKE 4 x * * = ee * x ae E xx eeXk 6 xe *xXKe we * 3 K at a eb ee ex S * Fi x . : : * * e . * * f * * Baseline sl 3 * Policy ol 10 20 30 Years After Policy Implementation Figure 7: The policy reduces the average number of drugs entering the market, but sub- stantial year-to-year variation occurs. In 10 simulations, after year 5 the number of new drugs per year with the policy (red dots) starts to diverge from the number of new drugs without the policy (X marks). Boldface symbols represent the average number across the simulations. Figure 7 presents average and individual year simulations. Uncertainty exists in the simulations because movement through the process is based on draws from distributions and individual success rates for each phase. The figure presents 10 of 100 simulations. The figure illustrates just how much uncertainty is generated naturally. Black X marks represent the number of new drugs entering the market under the baseline policy. The 23 X marks are spread out in a large cloud, indicating that the simulation model produces a large amount of variation in the number of drugs entering the market from year to year. Red dots represent the number of new drugs entering the market under the policy. Again, the cloud of red dots indicates a large amount of variation from the simulation. The figure shows that although on average the policy leads to fewer drugs entering the market, for any particular year the same, fewer, or more drugs could enter the market under the policy in comparison with the baseline. Whereas the policy tends to reduce the number of new drugs entering the market, the natural variation may lead to an increase in the number of new drugs entering the market in any particular year. In the middle two-thirds of the simulations (that is, between the 17th and 83rd percentiles), the reduction in the number of new drugs entering in the third decade after implementation of the policy ranges between 21 and 59. Because of uncertainty about the modeling framework itself, CBO expects that the range in which two-thirds of future outcomes would fall is wider than that from those simulations alone. 9 Conclusion This working paper describes a model CBO uses to inform its estimates of how various policies affect. development of new drugs. The model considers the firm's decision at the start of the various phases of human clinical trials. The firm considers expected cost and expected returns of entering the phase. The paper considers what happens when a policy is introduced that reduces the top quintile of expected returns by 15 percent to 25 percent. Using the model, CBO estimates that such a policy would reduce the number of drugs entering the market by 0.5 percent in the first decade under the policy. Owing to an accumulated effect through the phases, CBO estimates the number of drugs entering the market decreases by 8 percent in the third decade under the policy. The illustrative policy's exact implications for the health of families in the United States are unclear. CBO has estimated neither which types of drugs may be affected nor how the reduction in the number of new drugs will affect health outcomes. In addition, the policy may lead to lower prices and increased usage for drugs already on the market. CBO has not determined the overall effect of the policy on health outcomes. 10 Appendix: Accounting for Reduced Earnings In the preceding analysis, the Congressional Budget Office assumes that a policy such as the negotiation policy in H.R. 3 affects pharmaceutical development only through changes to the expected profitability of new investments at a fixed cost of financing. By reducing the earnings available to pharmaceutical firms to finance new development without tapping external sources, the policy conceivably could raise their cost of financing, further affecting drug development. Large drug companies are profitable enough now to finance R&D almost 24 Baseline EM ee MK ee KE MMR EK y Ry Number of Crugs Entering Market 25 30 Years After Policy Implementation (December 10). www. cho. gov/system/files/2019-12/hr3_complete.pdf (230 KB). Damodaran, Aswath. 2020. "Cost of Capital by Sector (U.S.)" (accessed December 8, 2020). http: //tinyurl.com/171f£0kqm. Department of State. 2010. "Framework for Promoting Transatlantic Economic Integration, Annex I: Fostering Cooperation and Reducing Regulatory Barriers, B. Sectoral Cooperation- Medicinal Products" (January 24). https://go.usa.gov/xAMHK. DiMasi, Joseph A., Ronald W. Hansen, and Henry G. Grabowski. 2003. "The Price of Innovation: New Estimates of Drug Development Costs." Journal of Heath Economics 22(2):151-185. https: //doi.org/10.1016/S0167-6296 (02)00126-1. DiMasi, Joseph A., Henry G. Grabowski, and John Vernon. 2004. "R&D Costs and Returns by Therapeutic Category." Drug Information Journal 38:211-223. https: //doi.org/10. 1177/009286150403800301. DiMasi, Joseph A. 2013. "Causes of Clinical Failures Vary Widely by Therapeutic, Phase of Study." Impact Report 15(5). Tufts Center for the Study of Drug Development. https: //csdd.tufts.edu/impact-reports. DiMasi, Joseph A., Henry G. Grabowski, and Ronald W. Hansen. 2016. "Innovation in the Pharmaceutical Industry: New Estimates of R&D Costs." Journal of Heath Economics 47:20-33. https: //doi.org/10.1016/j.jhealeco.2016.01.012. Dranove, David, Craig Garthwaite, and Manuel I. Hermosilla. 2020. Expected Profits and the Scientific Novelty of Innovation. Working Paper 20-16 (Northwestern University Insti- tute for Policy Research). http: //tinyurl.com/44smzy72. Dubois, Pierre, Olivier de Mouzon, Fiona Scott-Morton, and Paul Seabright. 2015. "Mar- ket Size and Pharmaceutical Innovation." RAND Journal of Economics 46(4):844-871. https://doi.org/10.1111/1756-2171.12113. Food and Drug Administration (FDA). 2019. "Development & Approval Process: Drugs." https://go.usa.gov/xAMsH. Hansen, Lars Peter. 1982. "Large Sample Properties of Generalized Method of Moments Estimators." Econometrica 50(4):1029-1054. https: //doi.org/10.2307/1912775. Heckman, James, and Bo Honoré. 1990. "The Empirical Content of the Roy Model." Econo- metrica 58(5):1121-1149. https: //doi.org/10.2307/2938303. Heckman, James J., and Edward J. Vytlacil. 2007. "Econometric Evaluation of Social Pro- grams, Part I: Causal Models, Structural Models, and Econometric Policy Evaluation," in James J. Heckman and Edward E. Leamer, eds., Handbook of Econometrics, volume 6B (Elsevier, 2007), pp. 4779-4874. https: //doi.org/10.1016/S1573-4412(07)06070-9. 27 Khmelnitskaya, Ekaterina. 2020. "Competition and Attrition in Drug Development." Uni- versity of Virginia. https: //tinyurl . com/306loghf. Wong, Chi Heem, Kien Wei Siah, and Andrew W. Lo. 2019. "Estimation of Clinical Trial Success Rates and Related Parameters." Biostatistics 20(2):273-286. https: //doi.org/ 10.1093/biostatistics/kxx069. 28