Data and Program README file for -Y´Trade and Market Selection: Evidence from Manufacturing Plants in Colombia¡ by Marcela Eslava, John Haltiwanger, Adriana Kugler, and Maurice Kugler. The empirical analysis in this paper uses proprietary establishment level data from a database constructed by the authors using information from the Colombian Annual Manufacturing Survey (AMS). The AMS is housed at the Colombian Departamento Administrativo Nacional de Estadísticas (DANE). The specific data for this paper, generated by the authors using information from the AMS, has been archived by the Dirección de Metodología y Producción Estadística at DANE. For approved projects and for statistical purposes only, the data for this project can be accessed through DANE. For access to the data, interested users must contact DANE following the instructions available at: http://www.dane.gov.co/files/investigaciones/industria/eam/pasos.html Two .zip files, each containing one folders, are included. The folder Baseline-programs contains the SAS and Stata programs necessary for generating the tables and figures in this paper (see below for detailed description). The folder Bases-PA contains SAS and Excel databases with public domain, or otherwise unrestricted, data used by the programs (see below for detailed description). For replication, these public domain databases and the proprietary data obtained from DANE must be saved to a directory: C:\Trabajo\Published Papers\Survival Rev\Backup Sent September 14, 2012\Public Acces\Bases-run\ Another directory must be created to store the log files (Stata log files will be automatically stored in this directory), called: C:\Trabajo\Published Papers\Survival Rev\Public acces\logs Versions of software used: Stata SE 11.2 for Windows, SAS 9.2. Programs: The programs necessary to replicate the results (some SAS programs, some STATA programs) are listed below in the order in which they must be processed: 3. survival_intro_PA.sas: feeds input databases and organizes data. 4. tables 1-2 & Herfindahl_PA.sas: generates information for Tables 1 and 2 (and other versions of the same descriptive statistics, for instance from different levels of aggregation). 5. table 3_PA.sas: Estimates demand function of Table 3 and stores demand shock and demand elasticity. Note: robust standard errors for this estimation not reported here, but in next program. 5b. demand std err_PA.do: This program estimates again demand function as in program 5 but generating robust standard errors and first stage R2s reported in the tables. 6. table 4_PA.sas: Organizes data for exit model estimation, generates numbers for Table 4. 7. Probit table WA3 no tariffs.do: Baseline probit model [Table 5] but without interactions with tariffs. Produces Web Appendix table WA3. 8. Probit table 5.do: Baseline probit model [Table 5]. Also runs the same model replacing GDP growth with time effects [Table 10]. Calculates marginal effects at different levels of tariffs and differences between marginal effects at 20% and 60% tariffs, with significance levels. 9. Transit to dynamo for elasticity.do: Creates databases for dynamic symultation 10. Tables 6-7.sas: Dynamic simulation of tables 6 and 7. Runs 1,000 iterations and collects average results. 11. Transit to table 8.sas: structures databases for Table 8 12. Table 8.do: Table 8 regressions 12a. table 8 for figure 2.do, saves weighted TFP information for Figure 2. 13. figures 1-2.sas: Generates Figures 1 and 2. Important: excel files Figure2 and Figure 3 should already exist in the -Y´Bases-run¡ folder and have a sheet with the figure, so that the data are only updated. This is why the attached databases include these files. 14. Probit table 10 K.do Runs exit probit controlling for capital stock for Table 10. 15. Probit table 11 Herf out.do: Runs exit probit not considering most concentrated sectors for Table 11. 16. Probit table 11 Change in tariffs .do Runs probit controlling for change in tariffs for Table 11. 17. Probit table 9 TFP factor shares (folder): Contains a series of programs that must be processed in the order implied by file names (from 4 to 8) to generate Table 9 estimates. Databases: Plant-level (available from DANE, proprietary data): Varbas.sas7bdat: plant-year is the unit of observation. Contains all plant-level variables at constant prices the paper uses. This dataset was initially constructed using plant level information from the AMS. For this paper, we simply take it directly from DANE. The database contains the following variables: VARIABLE NAME IN ‘VARBAS.sas7bdat’ DESCRIPTION Pro Output at constant prices K Capital (Buildings, structures, machinery and equipment) Direc Non-production Personnel perpr Production Personnel Labor Employees (Production and non-production personnel) Mat Materials at constant prices En Energy Consumption pr_eni Energy price index pr_prod Output price index pr_mat Materials price index Rpren Energy relative price (logarithmic difference between price index and PPI) rprmat Materials relative price (logarithmic difference between price index and PPI) Relpr Output relative price (logarithmic difference between price index and PPI) Year Ciiu3 3 digit sector code (ISIC revision 2) Ciiu4 4 digit sector code (ISIC revision 2) Id Plant ID (fake) Other variables are included in DANE’s varbas dataset, but are deleted by the first program 1. because they are not used for this paper, or they are re-built during the process (Dshsect, Dsh, Tfp1, and others). Aggregate or sector level, or unrestricted (included in this .zip file): ppi.sas7bdat: Producer Price Index. Source: Banco de la República (Banrep: www.banrep.gov.co/). PPI data, as reported by Banrep, are monthly. The annual series we present are the simple average for each year of the correspondent index, indexed to 1 in 1982. BDIReformas.xls: Lora reform index and its disaggregate categories. Source: IDB. Agnpool2.sas7bdat: cost shares at 3-digit level. Built from a dataset originally called prodfpool. The latter is equal to dataset Varbas, but contains the full set of establishments contained in EAM, including those with missing values in deflated production. The program to generate this database is also available from the authors. Building these cost shares from Varbas yields very similar results. The database Agnpool2.sas7bdat contains the following variables: VARIABLE NAME IN ‘AGNPOOL2.sas7bdat’ DESCRIPTION Shl Average labor cost share in total output She Average Energy cost share in total output Shm Average Materials cost share in total output Shk Average Capital cost share in total output, obtained as a residual Ciiu3 3 digit sector code (ISIC revision 2)l Agnpool2a.sas7bdat: similar to Agnpool2.sas7bdat but with cost shares calculated for the pool of sectors. Program also available from the authors, also originally processed on the full set of AMS data, but also similar results are obtained if it is processed on Varbas. prodorgtot.sas7bdat: IV estimates of factor elasticites, for the pool of data. Program also available from the authors, also originally processed on the full set of AMS data, but also similar results are obtained if it is processed on Varbas. promedioaranceles: Effective tariffs for each 4-digit sector in each year. 4-digit level tariffs are obtained as simple averages of product tariffs at the level of 10-digit NANDINA codes. Source for the product tariffs: National Planning Department. Variables description: VARIABLE NAME IN ‘promedioaranceles.sas7bdat’ DESCRIPTION Aran_efect Average effective tariff for the sector Year Year Ciiu 4 digit sector code (ISIC revision 2)l code_stay: plant-year level dataset containing unrestricted information not contained in database -Y΄varbas‘, with plants identified with the fictitious ID codes used in ΄varbas‘. It also contains ΄stay‘ a variable that allows us to indicate if the plant exited as late as one year after we see no report of deflated output for the plant (΄varbas‘ contains only information on observations for which deflated production can be constructed). Description of variables: VARIABLE NAME IN ‘code_stay.sas7bdat’ DESCRIPTION Code Fictitious plant ID in database Varbas Year Year Ciiu 4 digit sector code, ISIC revision 2 (also contained in varbas) Ciiu3 3 digit sector code, ISIC revision 2 (also contained in varbas) Stay Indicates whether the plant reported to DANE in that year even if no report of deflated production could be constructed (so the observation may not be in varbas). Pib_dane Real GDP. Source: DANE Gdp_growth Real GDP growth. Own calculations based on -Y΄Pib_dane‘) Roads Indicator of whether we have information for the observation on state roads. This is used to restrict the set of observations that enter the demand estimation, for consistency with other work by the authors (only a handful of observations are discarded). DANE general contact information: www.dane.gov.co Address: Carrera 59 No.26-70 Interior I - CAN. Phone number: (571) 5978300 Bogotá D.C., Colombia %G–%@ South América contacto@dane.gov.co