<table >
<tr>
<td colspan="2" style="height: 14px">
tdtext1
<a>hyperlinktext1<a/>
</td>
</tr>
<tr>
<td>
tdtext2
</td>
<td>
<span>spantext1</span>
</td>
</tr>
</table>
This is my sample text. How to write a regular expression in C# to get the matches for the innertext for td, span, hyperlinks.
I cringe every time I hear the words regex and HTML in the same sentence. I would suggest checking out the HtmlAgilityPack on CodePlex which is a very tolerant HTML parser that lets you use XPath queries against the parsed document. It’s much cleaner and the person that inherits your code will thank you!
EDIT
As per the comments below, here’s some examples of how to get the InnerText of those tags. Very simple.