pith. sign in

arxiv: q-bio/0310022 · v1 · pith:2W7DWU5Tnew · submitted 2003-10-16 · 🧬 q-bio.GN

A covariance kernel for proteins

classification 🧬 q-bio.GN
keywords kernelbiologicalclassificationproteinsequencesachievementsamino-acidassumptions
0
0 comments X
read the original abstract

We propose a new kernel for biological sequences which borrows ideas and techniques from information theory and data compression. This kernel can be used in combination with any kernel method, in particular Support Vector Machines for protein classification. By incorporating prior biological assumptions on the properties of amino-acid sequences and using a Bayesian averaging framework, we compute the value of this kernel in linear time and space, benefiting from previous achievements proposed in the field of universal coding. Encouraging classification results are reported on a standard protein homology detection experiment.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.