
TL;DR
This paper reveals a fundamental link between parsimony scores and tree rearrangements for unordered, reversible characters, impacting phylogenetic analysis and understanding of incongruence and recombination.
Contribution
It establishes that for unordered, reversible characters, the parsimony score equals the number of tree rearrangements needed, offering new insights into consensus trees and character incongruence.
Findings
Parsimony score equals the number of tree rearrangements for certain characters.
Implications for consensus tree debates and total evidence approaches.
Connection between character incongruence and recombination.
Abstract
The parsimony score of a character on a tree equals the number of state changes required to fit that character onto the tree. We show that for unordered, reversible characters this score equals the number of tree rearrangements required to fit the tree onto the character. We discuss implications of this connection for the debate over the use of consensus trees or total evidence, and show how it provides a link between incongruence of characters and recombination.
Click any figure to enlarge with its caption.
Figure 1
Figure 2Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Title:
Parsimony via consensus
Trevor C. Bruen and David Bryant
Abstract
The parsimony score of a character on a tree equals the number of state changes required to fit that character onto the tree. We show that for unordered, reversible characters this score equals the number of tree rearrangements required to fit the tree onto the character. We discuss implications of this connection for the debate over the use of consensus trees or total evidence, and show how it provides a link between incongruence of characters and recombination.
Running Head: Parsimony via consensus
Key Words: maximum parsimony, compatibility, supertree, matrix representation with parsimony, homoplasy
Corresponding Author: David Bryant
Introduction
The (Fitch) parsimony length of a character on a tree equals the minimum number of state changes (substitutions) required to fit the character onto a tree (Fitch, 1971). We turn this definition on its head and show how the parsimony length of a character equals the minimum number of changes in the tree required to fit the tree onto the character. This may be a back-to-front way to look at parsimony, but it is also a useful one. We detail two applications of the result.
The first application is that this reformulation of parsimony provides a closer link between parsimony based analysis and supertree methods. We demonstrate that the maximum parsimony tree can be viewed as a type of median consensus tree, where the median is computed with respect to the SPR distance (see below). As well, the result shows how to conduct a parsimony based analysis not just on characters but on trees, without having to recode the trees as binary character matrices. This opens the way to a hybrid between the consensus approach and the total evidence approach, where the data is a mix characters, trees, and subtrees.
The second application of our observation on parsimony is to the analysis of pairs of characters. We show that the score of the maximum parsimony tree for two characters is a simple function of the smallest number of recombinations required to explain the incongruence between the characters without homoplasy. This result provides the basis of a highly efficient test for recombination (Bruen et al., 2006).
Here and throughout the paper we assume that all phylogenetic trees are fully resolved (bifurcating) and that by ‘parsimony’ we refer to Fitch parsimony, where the character states are unordered and reversible. Some of the results presented here can be extended to other forms of parsimony, and possibly to incompletely resolved trees (Bruen, 2006), lie beyond the scope of this paper.
Note that in this paper we are dealing with unrooted SPR rearrangements, which are those used in tree searches. There is a related, but distinct, concept of rooted SPR rearrangements, where the rearrangements are restricted to obey a type of temporal constraint Song (2003). It is this latter class of rooted SPR rearrangments that are used to model lateral gene transfers and recombination. It would be a worthwhile, but challenging, goal to investigate whether any of the results on unrooted SPR rearrangements in this paper can be extended to rooted SPR rearrangements.
Linking Parsimony with SPR
A subtree-prune and regraft (SPR) rearrangement is an operation on phylogenetic trees whereby a subtree is removed from one part of the tree and regrafted to another part of the tree, see Figure 1, (Felsenstein, 2004; Swofford et al., 1996). These SPR rearrangements are widely used by tree searching software packages like PAUP (Swofford, 1998) and Garli (Zwickl, 2006). The SPR distance between two trees can be defined as the minimal number of SPR rearrangements required to transform one tree into the other (Hein, 1990; Allen and Steel, 2001; Goloboff, 2007). For example, the two trees and in Figure 1 can be transformed into each other using a minimum of two SPR rearrangements, via the tree , so their SPR distance is two.
The parsimony length of a character on a tree is the minimum number of steps required to fit that character on the tree, as computed by the algorithm of (Fitch, 1971). We will always assume unordered reversible characters The length of a character on a tree is denoted . A character with states therefore has parsimony length at least , as every state not at the root has to arise at least once. A character is compatible with a tree if it requires at most changes on that tree (Felsenstein, 2004).
So far, one thinks of fitting a character onto a tree; we could just as well fit the tree onto the character. If the character and the tree are compatible then we have a perfect fit. When there is not a perfect fit we can measure how many SPR rearrangements are required to give a tree that does make a perfect fit. It turns out that this measure gives an equivalent score to parsimony length. More formally:
Theorem 1**.**
Let be a character with states and let be a fully resolved phylogenetic tree. It takes exactly SPR rearrangements to transform into a tree compatible with . The result still holds if has some missing states.
As an example, consider the character mapping taxa A,C,D,F to one and B,E,G to zero. The length of this character on tree of Figure 1 is three, and the number of SPR rearrangements needed to transform onto some tree compatible with with is two. Note that there could be other trees compatible with are are further than two SPR rearrangements away: the result only gives the number of rearrangements required to obtain the closest tree.
Once stated, the theorem is not too difficult to prove. First show that performing an SPR rearrangement decreases the length by at most one step. Hence it takes at least SPR rearrangements to transform into a tree compatible with the character . Then show that this is the minimum required. A formal proof is presented in the Appendix. A restricted (binary character) version of this theorem was proved in (Bryant, 2003).
The theorem captures an issue that is central to the interpretation of incongruence: is an observed incongruence to be explained by positing homoplasy or by modifying the tree. Define the SPR distance from a tree to a character to be the SPR distance from to the closest tree that is compatible with . Theorem 1 then tells us that the SPR distance from is equal to the difference between the length of on and the minimum possible length of on any tree.
Consensus trees, supertrees and parsimony
In their insightful overview of supertree methods Thorley and Wilkinson (2003) characterise a family of supertree methods that all minimise a sum of the form
[TABLE]
Here are the input trees and is a measure of the distance between the input tree with the supertree . There are many choice for the distance measure , and it need not be the case that the distance measure satisfies the symmetry condition . Gordon (1986) was the first to propose this description of supertrees. Many supertree methods can be described in these terms, including Matrix representation with parsimony (MRP) (Baum, 1992; Ragan, 1992); Minimum Flip supertrees Chen et al. (2006); the Median Supertree (Bryant, 1997), Majority Rule Supertree (Cotton and Wilkinson, 2007) and the Average Consensus Supertree (Lapointe and Cucumel, 1997).
Let denote the SPR distance from to the closest fully resolved tree that is compatible with . By Theorem 1, a maximum parsimony tree for is one that minimises the expression
[TABLE]
In this way, maximum parsimony is a form of median consensus. The significance of this observation doesn’t come from the fact that we can write the the parsimony score of in the form (2); it is from the close connection with SPR distances, and from the way we will now use this connection to combine different kinds of data in the same theoretical framework.
An SPR median tree for fully resolved trees on the same leaf set is a tree that minimises
[TABLE]
where here denotes the SPR distance from to (Hill, 2007). We extend this directly to a supertree method by mimicking the situation for characters. Suppose that is a phylogenetic tree, not necessarily fully resolved, on a subset of the set of leaves. We say that a fully resolved tree on the full set of leaves is compatible with (equivalently, displays ) if we can obtain from by pruning off leaves and contracting edges. In this general situation, we let denote the SPR distance from to the closest fully resolved tree that is compatible with . This is equivalent to the more traditional definition whereby we first prune leaves off then compute the distance from this pruned tree to .
Now suppose that we have both characters and trees in the input. Both types of phylogenetic data can be into an SPR median tree , chosen to minimise the sum
[TABLE]
We have, then, a way to bring together both the supertree/consensus methodology and the total evidence methodology. In the case that the data comprises only trees, the tree is a median supertree; in the case that the data comprises only character data, the tree is the maximum parsimony tree.
It is important to note the difference between this approach and the MRP method (Baum, 1992; Ragan, 1992), which could be used to combine trees and characters. In MRP, the trees are broken down into multiple independent characters. This is a problem, since the characters encoding a tree are nowhere near independent. In contrast, the SPR median tree approach treats a tree as a single indivisible unit of information.
There is one critical issue that has been side-stepped: computation time. At present, computational limitations make the construction of SPR median trees infeasible for all but the smallest data sets: just computing the SPR distance between two trees is an NP-hard problem (Hickey et al., 2006). In contrast, Total evidence and MRP approaches are possible for at least 100 taxa. However there are now good heuristics for unrooted SPR distance Goloboff (2007) and exact special case algorithms Hickey et al. (2006) that could be applied to the problem. Below we describe a lower bound method for the SPR distance that should also aid construction of these SPR median trees.
Parsimony on pairs of characters
Another valuable application of Theorem 1 follows when we consider parsimony analysis of just two unordered and reversible characters. The concept of pairwise character compatibility was introduced by Le Quesne (1969) (see also Felsenstein (2004)). Two binary characters with states [math] and are incompatible if and only if all four combinations of , and are present as combination of states for the two characters (Le Quesne, 1969). In a standard setting, character incompatibility is interpreted as implying that at least one of the characters has undergone convergent or recurrent mutation (homoplasy). In other words, for every possible phylogeny describing the history of the two characters, at least one homoplasy is posited for one of the characters. Another interpretation of incompatibility of two characters is that characters evolved without homoplasy on two different phylogenies, where the phylogenies differ by one or more SPR rearrangement (Sneath et al., 1975; Hudson and Kaplan, 1985).
Define the total incongruence score for two multi-state unordered characters and (with and states respectively) as
[TABLE]
This is the maximum parsimony score of the two characters minus the minimum number of changes required for each character. Equation (3) generalises the incompatibility notion for two binary characters. It is also equivalent to the incongruence length difference statistic applied to only two characters (Farris et al., 1995). Importantly, the total incongruence score can be computed rapidly (Bruen and Bryant, 2006). The following consequence of Theorem 1 strengthens the connection between incongruence and SPR rearrangements.
Theorem 2**.**
The total incongruence score for two characters equals the minimum SPR distance between a tree and such that is compatible with and is compatible with .
Although the notion of total incongruence for two characters has been considered before in the context of character selection and weighting (Penny and Hendy, 1986), it has not been considered in the context of genealogical similarity. Essentially, Theorem 2 shows that the total incongruence score equals the minimum possible number of SPR rearrangements that could have occurred between the phylogenetic histories for both characters, assuming that the characters have different histories with which they are each perfectly compatible.
Indeed, Theorem 2 suggests a natural way to interpret genealogical similarity between two characters, which we have used to develop a powerful test for recombination (Bruen et al., 2006). Choosing two characters from two different genes (which have possibly different histories) gives a simple approach to identify the distinctiveness of the histories of the genes.
We can also apply Theorem 2 to obtain a lower bound on an SPR distance between two trees. Suppose that we have two trees and and we wish to obtain a lower bound on the SPR distance between the two trees. If we choose any character convex on and any character convex on then, by Theorem 2, we have that . By carefully choosing and we can obtain tighter bounds. One natural starting point for and is the four or five character encodings described by (Semple and Steel, 2002; Huber et al., 2005).
Discussion and extensions
We have presented a reformulation of parsimony that is, in some way, dual to the standard definitions. Instead of measuring how well a character fits onto a tree we look at how well the tree fits onto the character. A consequence of this new perspective is that we can combine trees and character data using one general SPR framework, and we also obtain new results connecting incongruence measures and recombination. Nevertheless, it is not immediately clear how the new reformulation can be interpreted in itself.
One aid in this direction is to consider the information a single character, or tree, represents. Given a single character, we can imagine a cloud of trees comprising exactly those trees compatible with the character (Figure 2). If we are told that this character evolved without homoplasy, then we know that the true evolutionary tree must be contained somewhere within the cloud. However as there is only one character there is a lot of uncertainty regarding the tree, so there are a lot of trees in the clouds. Now suppose we have multiple characters, each with its own cloud. There may not be a single tree contained in the intersection of all of these clouds. Instead, we search for a tree that is close as possible to all of the clouds. The distance from to the cloud associated to character is exactly , so by Theorem 1 a tree closest to all of the clouds is a maximum parsimony tree.
Each cloud represents the uncertainty around each piece of data (tree or character).
We note that several of the results in this article can be extended, for details. Firstly, both Theorems 1 and 2 are both valid if we replace the SPR distance with the tree bisection and reconnection (TBR) distance. In a TBR rearrangement, a subtree is removed from the tree and then reattached elsewhere in a tree, the difference with SPR being that we can reattach using any of the nodes in the subtree (Allen and Steel, 2001; Felsenstein, 2004). The TBR distance between two trees is the minimum number of TBR rearrangements required to transform one tree into the other.
That Theorems 1 and 2 hold for the TBR distance might seem surprising, since the TBR distance between two trees is always less than, or equal to, the SPR distance between the trees. However the extension follows by a tiny change to the proof of Theorem 1, noting that a TBR move can still only reduce the parsimony score of a character by at most one.
We have also explored extensions of the result to other distances between trees, notably the Robinson-Foulds or partition distance and the Nearest Neighbor Interchange distance, though the connections are not so clear. See Bruen (2006) for details.
Acknowledgements
We would like to thank Mike Steel, Sebastien Böcker, Olaf Bininda Emonds, Pablo Golloboff, Mark Wilkinson and an anonymous referee for their valuable suggestions. This research was partially supported by the New Zealand Marsden Fund.
Appendix
Refer to (Semple and Steel, 2003) for a detailed description of the notation.
The first observation is that an TBR rearrangement of a tree increases the length of a character by at most one. As SPR rearrangements are a special case of TBR rearrangements, the same result holds for SPR.
Lemma 1**.**
Let be a fully resolved phylogenetic tree and an unordered reversible character. Let be a phylogenetic tree that differs from by a single TBR rearrangement. Then .
Proof.
The proof of Lemma 5.1 in (Bryant, 2004) for binary characters applies directly to the multistate case. ∎
Let denote the unrooted SPR distance between two phylogenetic trees and .
Theorem 1 Let be a character with states and let be a fully resolved phylogenetic tree. It takes exactly SPR rearrangements to transform into a tree compatible with . The result still holds if has some missing states.
Proof.
Let be a fully resolved phylogenetic tree compatible with for which is minimized and let . Then there exists a sequence of trees such that every adjacent pair of trees in the sequence differ by exactly one SPR rearrangement. By Lemma 1 the existence of this sequence implies that and since is compatible with we have , giving
[TABLE]
For the other direction, we show that we can construct a sequence of SPR rearrangements that transform into a tree compatible with . Firstly, if , then is compatible with so the proof is finished. Otherwise, let be an assignment of states to internal nodes that minimises the number of state changes (that is, a minimum extension of ). Then since is not convex on there exist three vertices and , where , lies on the path from to and . Perform an SPR rearrangement by removing edge , supressing the vertex and creating a new edge where is a new vertex on an edge adjacent to . Furthermore, set . Then the number of edges on which a change has occurred has decreased by thereby decreasing the parsimony length by . This procedure can be repeated until the parsimony length equals , constructing the desired sequence of trees and completing the proof. ∎
Let be a maximum parsimony phylogenetic tree for and and let
Theorem 2 The total incongruence score for two characters equals the minimum SPR distance between a tree and such that is compatible with and is compatible with .
Proof.
Let and be any two trees compatible with and respectively. Then and by Theorem 1, . We have then that
[TABLE]
and so is a lower bound for .
We show that this bound can be achieved. Let be a maximum parsimony tree for the pair of characters . By Theorem 1 there exist two trees and such that is compatible with , is compatible with and
[TABLE]
implying that and hence
[TABLE]
∎
The reference list from the paper itself. Each links out to its DOI / PubMed record.
- 1Allen and Steel (2001) Allen, B. and M. Steel. 2001. Subtree transfer operations and their induced metrics on evolutionary trees. Annals of Combinatorics 5:1–13.
- 2Baum (1992) Baum, B. 1992. Combining trees as a way of combining datasets for phylogenetic inference, and the desirability of combining gene trees. Taxon 41:3–10.
- 3Bruen (2006) Bruen, T. 2006. Discrete and statistical approaches to genetics. Ph.D. thesis Mc Gill University School of Computer Science.
- 4Bruen and Bryant (2006) Bruen, T. and D. Bryant. 2006. A subdivision approach to maximum parsimony. Annals of Combinatorics In Press.
- 5Bruen et al. (2006) Bruen, T., H. Philippe, and D. Bryant. 2006. A simple and robust statistical test to detect the presence of recombination. Genetics 172:1–17.
- 6Bryant (1997) Bryant, D. 1997. Building trees, hunting for trees and comparing trees. Ph.D. thesis Dept. Mathematics, University of Canterbury.
- 7Bryant (2003) Bryant, D. 2003. A classification of consensus methods for phylogenetics. Pages 163–184 in Bioconsensus vol. 61 of DIMACS . American Math Society, Providence, RI.
- 8Bryant (2004) Bryant, D. 2004. The splits in the neighborhood of a tree. Annals of Combinatorics 8:1–11.
