Supplementary: Inference of Spatial Organizations of Chromosomes Using Semi-‐definite Embedding Approach and Hi-‐C Data Zhizhuo Zhang1 Guoliang Li2 Kim-‐Chuan Toh3 1
Wing-‐Kin Sung1,2 3
School of Computing, National University of Singapore & Department of Mathematics, National University of 2
Singapore & Genome Institute of Singapore 1
2
3
{Zhizhuo,ksung}@comp.nus.edu.sg & [email protected]‐star.edu.sg & [email protected] Glossary 3D: three dimension SDP: semi-‐definite programming Data Preparation: Simulation Data: The simulation data was generated from three different types of manually-‐created 3D structures (Supp Figure 4): (1)Helix curve, (2)Brownian motion simulation of a single particle and (3)Uniform random points in a cube. Each structure is represented by 100 points. We assume that the Hi-‐C technique is sensitive enough to capture interactions with at most 50 nearest neighbours and the conversion factor α is 1, i.e., the interaction frequency f of two given points can be computed as f=(1/d)1/ =1/d (d is the spatial distance between given points). Further, we scale the frequency matrix so that the summation of frequency equals to 100000, which is similar to real α
Hi-‐C data and help to get better performance for BACH (not our ChromSDE) program that is sensitive to the scale of data. Real Data mESC Hi-‐C pair-‐read mapped data were generated by Dixon JR et al.[1] and were downloaded from NCBI with accession number GSE35156. The mapped data were further normalized using the normalization pipeline developed by Yaffe E. and Tanay A.[2] which produce the normalized contact frequency matrix and raw contact frequency matrix for different resolution. GM Hi-‐C normalized frequency matrix and raw frequency matrix were obtained directly from [2], which were downloaded from Tanay group website: http://compgenomics.weizmann.ac.il/tanay/?page_id=283 mESC Pol2 ChIA-‐PET data is generated by our lab and will be published with another paper later. Tested Programs’ Settings: All tools were tuned to their best performances by our best effort in all experimental settings. All programs tested in this paper have standalone version, and they were tested in our Linux server with 144GB RAM and two Intel X5670 CPUs. BACH and MCMC5C are the only two programs we can find which are capable for constructing 3D chromatin structures from general Hi-‐C data. Other 3D modeling programs mentioned in the publications[3, 4] are either too specific or not available. For example, IPOPT based program published in [3] convert the 4C frequency to spatial distance by hard-‐coded condition which is not capable for different scale or sequencing depth data. The author[5] who implemented 3D chromosome modeling program based Integrated Modeling Platform (IMP) said the original program was not workable anymore due to the platform upgrade. Below is the setting of all tested programs in the paper. ChromSDE: Available from http://biogpu.ddns.comp.nus.edu.sg/~zzz/ChromSDE/ The script is executed in Matlab 2009. Linear SDP applies Yalmip [6] as modeling language which solves standard SDP problem by a general SDP solver called SDPT3 [7]. “dualize” option in Yalmip is used to speed up when the input distance matrix is sparse. Quadratic SDP applies a specialized solver called PPApack developed by Toh K.C.[8]. If not specially mentioned, the result of ChromSDE is generated by Quadratic SDP. BACH: Available from http://www.people.fas.harvard.edu/~junliu/BACH/ Command: BACH -‐i matrixfile -‐v featurefile -‐K 100 -‐MP 10 -‐NG 5000 -‐NT 50 -‐L 50 -‐SEED 1 Comment: maxtrixfile is the raw contact frequency matrix file, and featurefile is the information of CG content, fragment length and mappability for each genomic region. The enzyme feature information is generated using the script from the same group , which were downloaded from: http://www.people.fas.harvard.edu/~junliu/HiCNorm/ BACH*: Command: BACH -‐i normalized_matrixfile -‐v fake_featurefile -‐K 100 -‐MP 10 -‐NG 5000 -‐NT 50 -‐L 50 -‐SEED 1 Comment: normalized_matrixfile is the normalized contact frequency matrix file, and fake_featurefile is the file contained the fake enzyme feature information. We assigned CG content, fragment length and mappability as a random value uniformly sampled from (0,1), and ensures the BACH program does not take these information into account for explaining the input frequency. The program cannot accept constant value for those features, and the current way of handling the normalized data is suggested by the author of BACH. MCMC5C: Available from http://dostielab.biochem.mcgill.ca/tools.php Command: (simulation data) java -‐jar MCMC5C_Galaxy.jar IFfile FragmentFile 100000000 100 0.05 1 6 dump.txt 0.80 0.10 0.10 -‐50.0 (real data) java -‐jar MCMC5C_Galaxy.jar IFfile FragmentFile 100000000 100 0.05 2 6 dump.txt 0.80 0.10 0.10 -‐50.0 Comment: For simulation data, the sixth parameter is 1 which corresponds to the conversion factor equal to 1 in our experiment setting. For real data, the sixth parameter is 2 which corresponds to the conversion factor equal to 0.5 and is suggested as their default setting. IFfile is the file containing the normalized contact frequency matrix and the prior variances of each given frequency. The prior variance is defined as sqrt(frequency+10) in their paper[9]. FragmentFile is the list of genomic loci presented by the matrix. Supplementary Figures: Supp Figure 1: Simulation result that compares the performance of different regularization parameters. The results were generated using simulation on a Brownian Motion Curve as stated in Section 3.1. In the experiments, different regularization parameter lambda (0.1, 0.01, 0.001) and different types of SDPs (linear, quadratic) combination were tested. (a) Spearman correlation between the pair-‐wise distance matrices of the predicted structure and the true structure under different noise level. (b) The absolute error of the estimated value of conversion factor under different noise levels. Supp Figure 2: The effect of different conversion factors on the 3D structures predicted by ChromSDE. The true value of conversion factor is 1. Each sub-‐figure is the predicted structure by ChromSDE with the specific conversion factor indicated below it. Supp Figure 3 Absolution error of estimated frequency with different values of conversion factor. (a)Simulation study: the data is generated by helix curve under different noise level and true conversion factor is 1. (b)Real Hi-‐C data: the data is from chromosome 16-‐18 of mESC Hind3 dataset. Supp Figure 4: Different types of structures used in the simulation study: Helix curve(Left), Brownian motion simulation of a single particle(Middle) and Uniform random points in a cube(Right). Supp Figure 5: Simulation result that compares the performance of different methods. (a)Root mean square deviation (RMSD) between the predicted structure and true structure(Brownian curve). Generally, RMSD increases as the noise level increases. The linear SDP and quadratic SDP of ChromSDE performs similarly below noise level 0.7 and better than other methods. (b) Consensus Index predicted by two SDP formulations under different noise levels. Generally, the consensus index decreases as the noise level increases. Two SDP formulations get quite similar values of consensus index. (c) The value of each bar presents the average conversion factor of different mix factors in Figure 3(d) for given noise level and the error bar represents the standard deviation of estimated conversion factor. The result shows the estimated conversion factors are very close to the true value 1 under different condition. (d) Consensus index decreases when the proportion of the dominate structure in the mixture decreases or noise level increases. For dominate structure ratio 1 to 0.5, the results were generated by simulating two Brownian structures with mix factor 1 to 0.5 correspondingly. For dominate structure ration 0.333, 0.2 and 0.1, the results were generated by simulating a mixture with 3, 5 and 10 Brownian structures with equal proportion correspondingly. (a) (b) Supp Figure 6: 3D structures predicted by ChromSDE using different enzyme data (red: Hind3, green : NcoI). The 3D structures are built using 1Mbp resolution data, and quadratic SDP . (a) mouse ES cell. (b) human GM cell. Supp Table 1: The conversion factors estimated by ChromSDP and BACH. Each table element is the mean value of the estimated conversion factor across all chromosomes, and the value in each bracket is the standard deviation of the corresponding mean. Conversion Factor Estimation Quadratic SDP Linear SDP BACH BACH* mESC_NcoI 0.5455(0.0167) 0.5437(0.0153) 0.4130(0.0129) 0.4285(0.0439) mESC_Hind3 0.5354(0.0145) 0.5390(0.0183) 0.4182(0.0143) 0.4408(0.0902) GM_NcoI 0.6284(0.0489) 0.5780(0.0382) 0.5942(0.0487) 0.7078(0.3104) GM_Hind3 0.6381(0.0568) 0.6075(0.0490) 0.7342(0.3487) 0.8818(0.6990) Supp Figure 7. The consensus indices for different Hi-‐C datasets. (a) The consensus indices estimated across different chromosomes using Hind3 enzyme (blue) and NcoI enzyme(red) in mESC Hi-‐C datasets. (b) The consensus indices estimated across different chromosomes using Hind3 enzyme (blue) and NcoI enzyme(red) in GM Hi-‐C datasets. Supp Figure 8. The performance of existing methods under data of different resolutions. The conversion factor of MCMC5C is set to be 0.5 for different resolution, and its predicted structures under different resolutions are not so similar. The conversion factors predicted by BACH and BACH* increase when the resolution increases. Reference: 1. Dixon, J.R., et al., Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature, 2012. 485(7398): p. 376-‐80. 2. Yaffe, E. and A. Tanay, Probabilistic modeling of Hi-‐C contact maps eliminates systematic biases to characterize global chromosomal architecture. Nature genetics, 2011. 43(11): p. 1059-‐65. 3. Duan, Z., et al., A three-‐dimensional model of the yeast genome. Nature, 2010. 465(7296): p. 363-‐7. 4. Russel, D., et al., Putting the pieces together: integrative modeling platform software for structure determination of macromolecular assemblies. PLoS biology, 2012. 10(1): p. e1001244. 5. Marti-‐Renom, M.A. and L.A. Mirny, Bridging the resolution gap in structural modeling of 3D genome organization. PLoS computational biology, 2011. 7(7): p. e1002125. 6. Lofberg, J. YALMIP: A toolbox for modeling and optimization in MATLAB. 2004. IEEE. 7. Toh, K.C., M.J. Todd, and R.H. Tütüncü, SDPT3—a MATLAB software package for semidefinite programming, version 1.3. Optimization Methods and Software, 1999. 11(1-‐4): p. 545-‐581. 8. Jiang, K.F.a.S., D.F. and Toh, K.C., A partial proximal point algorithm for nuclear norm regularized matrix least squares problems. National University of Singapore, 2012. preprint. 9. Rousseau, M., et al., Three-‐dimensional modeling of chromatin structure from interaction frequency data using Markov chain Monte Carlo sampling. BMC bioinformatics, 2011. 12(1): p. 414.
© Copyright 2026 Paperzz