How to write a regular expression for html parsing?

Question

I'm trying to write a regular expression for my html parser.

I want to match a html tag with given attribute (eg. <div> with class="tab news selected" ) that contains one or more <a href> tags. The regexp should match the entire tag (from <div> to </div>). I always seem to get "memory exhausted" errors - my program probably takes every tag it can find as a matching one.

I'm using boost regex libraries.

Beware of Zalgo
– Kelly S. French
Commented Jan 12, 2012 at 22:47 — Kelly S. French, Commented Jan 12, 2012 at 22:47

Community · Accepted Answer · 2017-05-23 12:14:16Z

7

You should probably look at this question re. regexps and HTML. The gist is that using regular expressions to parse HTML is not by any means an ideal solution.

edited May 23, 2017 at 12:14

CommunityBot

11 silver badge

answered Apr 27, 2009 at 8:46

Brian Agnew

273k38 gold badges341 silver badges443 bronze badges

Add a comment |

Community · Accepted Answer · 2017-05-23 12:28:34Z

2

You may also find these questions helpful:

Can you provide some examples of why it is hard to parse XML and HTML with a regex?

Can you provide an example of parsing HTML with your favorite parser?

edited May 23, 2017 at 12:28

CommunityBot

11 silver badge

answered Apr 27, 2009 at 20:26

Chas. Owens

65.1k24 gold badges139 silver badges232 bronze badges

Add a comment |

anonanon · Accepted Answer · 2009-04-27 08:53:23Z

2

As others have said, don't use regexes if at all possible. If your code is actually XHTML (i.e. it is also well-formed XML) aI can recommend both the Xerces and Expat XML parsers, which will do a much betterv job for you than regexes.

answered Apr 27, 2009 at 8:53

anon

Add a comment |

zajcevzajcev · Accepted Answer · 2009-04-27 13:08:14Z

1

Maybe regexps aren't the best solution, but I'm already using like five different libraries and boost does fine when it comes to locating <a href> tags and keywords.

I'm using these regexps:

/<a[^\n]*/searched attribute/[^\n]*>[^\n]*</a>/ for locating <a href> tags and:

/<a[^\n]*href[[^\n]*>/searched keyword/</a>/ for locating links

(BTW can it be done better? - I suck at regex ;))

What I need now is locating tags containing <a href>'s and I think regexps will do all right - maybe I'll need to write my own parsing function as piotr said.

answered Apr 27, 2009 at 13:08

zajcev

It's not that regular expressions are not the best solution - for what you're trying to do regex is not a valid solution at all. Use a HTML or XML parser instead.
– Peter Boughton
Commented Apr 27, 2009 at 13:21
Ok, so which one do you recommend. I'd prefer an easy one ;)
– zajcev
Commented Apr 27, 2009 at 16:41

Add a comment |

piotr · Accepted Answer · 2009-04-27 10:44:35Z

0

Do as flex does: match <div> with a case insensitive match, and put your parser in a "div matched" state, keep processing input until </div> and reset state.

This takes two regexps and a state variable.

SGML tags valid characters are [A-Za-z_:]

So: /<[A-Za-z_:]+>/ matches a tag.

answered Apr 27, 2009 at 10:44

piotr

5,7952 gold badges39 silver badges63 bronze badges

Or, instead of re-inventing the wheel, use an existing parser which has already been written and will already deal with edge cases and so on.
– Peter Boughton
Commented Apr 27, 2009 at 13:15

Add a comment |

Collectives™ on Stack Overflow

How to write a regular expression for html parsing?

5 Answers 5

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

5 Answers 5

Your Answer

Sign up or log in

Post as a guest

Linked

Related