2

给定以下 HTML:

<p><span class="xn-location">OAK RIDGE, N.J.</span>, <span class="xn-chron">March 16, 2011</span> /PRNewswire/ -- Lakeland Bancorp, Inc. (Nasdaq:   <a href='http://studio-5.financialcontent.com/prnews?Page=Quote&Ticker=LBAI' target='_blank' title='LBAI'> LBAI</a>), the holding company for Lakeland Bank, today announced that it redeemed <span class="xn-money">$20 million</span> of the Company's outstanding <span class="xn-money">$39 million</span> in Fixed Rate Cumulative Perpetual Preferred Stock, Series A that was issued to the U.S. Department of the Treasury under the Capital Purchase Program on <span class="xn-chron">February 6, 2009</span>, thereby reducing Treasury's investment in the Preferred Stock to <span class="xn-money">$19 million</span>. The Company paid approximately <span class="xn-money">$20.1 million</span> to the Treasury to repurchase the Preferred Stock, which included payment for accrued and unpaid dividends for the shares. &#160;This second repayment, or redemption, of Preferred Stock will result in annualized savings of <span class="xn-money">$1.2 million</span> due to the elimination of the associated preferred dividends and related discount accretion. &#160;A one-time, non-cash charge of <span class="xn-money">$745 thousand</span> will be incurred in the first quarter of 2011 due to the acceleration of the Preferred Stock discount accretion. &#160;The warrant previously issued to the Treasury to purchase 997,049 shares of common stock at an exercise price of <span class="xn-money">$8.88</span>, adjusted for stock dividends and subject to further anti-dilution adjustments, will remain outstanding.</p>

我想获取<span>元素内的值。我还想获取元素class属性的值。<span>

理想情况下,我可以通过一个函数运行一些 HTML,然后取回一个提取实体的字典(基于<span>上面定义的解析)。

上面的代码是一个较大的 HTML 源文件的片段,它无法与 XML 解析器相匹配。所以我正在寻找一个可能的正则表达式来帮助提取感兴趣的信息。

4

3 回答 3

9

使用此工具(免费): http ://www.radsoftware.com.au/regexdesigner/

使用这个正则表达式:

"<span[^>]*>(.*?)</span>"

第 1 组(每个匹配项)中的值将是您需要的文本。

在 C# 中,它看起来像:

            Regex regex = new Regex("<span[^>]*>(.*?)</span>");
            string toMatch = "<span class=\"ajjsjs\">Some text</span>";
            if (regex.IsMatch(toMatch))
            {
                MatchCollection collection = regex.Matches(toMatch);
                foreach (Match m in collection)
                {
                    string val = m.Groups[1].Value;
                    //Do something with the value
                }
            }

修改为回答评论:

            Regex regex = new Regex("<span class=\"(.*?)\">(.*?)</span>");
            string toMatch = "<span class=\"ajjsjs\">Some text</span>";
            if (regex.IsMatch(toMatch))
            {
                MatchCollection collection = regex.Matches(toMatch);
                foreach (Match m in collection)
                {
                    string class = m.Groups[1].Value;
                    string val = m.Groups[2].Value;
                    //Do something with the class and value
                }
            }
于 2011-03-16T15:53:22.337 回答
2

假设您没有嵌套的跨度标签,以下应该可以工作:

/<span(?:[^>]+class=\"(.*?)\"[^>]*)?>(.*?)<\/span>/

我只对其进行了一些基本测试,但它会匹配 span 标签的类(如果存在)以及数据,直到标签关闭。

于 2011-03-16T15:39:35.917 回答
1

强烈建议您为此使用真正的 HTML 或 XML 解析器。您无法使用正则表达式可靠地解析 HTML 或 XML ——您能做的最多就是接近,并且越接近,您的正则表达式就越复杂和耗时。如果你有一个大的 HTML 文件要解析,它很可能会破坏任何简单的正则表达式模式。

正则表达式之类<span[^>]*>(.*?)</span>的将适用于您的示例,但是有很多 XML 有效的代码很难甚至不可能用正则表达式解析(例如,<span>foo <span>bar</span></span>会破坏上述模式)。如果您想要一些可以在其他 HTML 示例上工作的东西,那么 regex 不是这里的方法。

由于您的 HTML 代码不是 XML 有效的,请考虑HTML Agility Pack,我听说它非常好。

于 2011-03-16T15:53:18.160 回答