• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
Discriminative Structured Models for Biological Sequence Analysis.

Discriminative Structured Models for Biological Sequence Analysis.

Paperback

General Mathematics

Currently unavailable to order

ISBN10: 1243662808
ISBN13: 9781243662804
Publisher: Proquest Umi Dissertation Pub
Pages: 210
Weight: 0.85
Height: 0.44 Width: 7.44 Depth: 9.69
Language: English
Building predictive models in computational biology involves three key elements: choosing an appropriate scoring model for the task at hand; developing efficient inference algorithms for making predictions; and optimizing scoring parameters so that, the generated predictions are biologically meaningful. In many sub areas of computational biology, research in predictive methods for biology has focused on the first, two steps, whereas methods used for scoring parameter estimation often rely on a scattered combination of techniques, ranging from ad hoc statistical analysis and physicochemical arguments to manual trial-and-error. In this thesis, we consider the problem of scoring parameter estimation for three key problems in computational biology: protein sequence alignment, RNA secondary structure prediction, and RNA simultaneous folding and alignment. We formulate the model estimation task as a special class of supervised machine learning problems where the goal is to learn a mapping from a structured input space (e.g., amino acid or RNA sequences) to a structured output space (e.g., alignments or foldings). Under this framework, the problem of model estimation reduces to solving a convex optimization problem. Following this setup, we design structured probabilistic or max-margin models for each task. To allow our algorithms to scale efficiently to large-scale training sets, we develop new fast online and batch convex optimization algorithms specially tailored for learning structured models. We also develop an automated approach for designing custom regularization penalties to prevent overfitting in feature-rich scoring models. The resulting software packages for alignment (CONTRAlign), secondary structure prediction (CONTRAfold), and simultaneous alignment and folding (RAF) each obtain state-of-the-art accuracy in their respective domains. In particular, our alignment algorithm, CONTRAlign, obtains substantially improved sensitivity for the difficult class of twilight zone alignments. Our RNA secondary structure prediction algorithm, CONTRAfold, achieves higher general accuracy than existing classical methods, demonstrating for the first time that a statistically estimated scoring model can outperform thermodynamic approaches. Finally, our RNA simultaneous folding and alignment program, RAF, achieves high accuracies while also taking advantage of new sparsity heuristics to achieve running times orders of magnitude faster than previous approaches.

Also in

General Mathematics