comment on

Hi Eric. You might consider looking into Andrei Mikheev's article on text segmentation in Handbook of Computational Linguistics and the chapter on parsing in the same book.

If you can give me some concrete examples of what you are looking to do, I might be able to scare up some info for you. I have to say that regular expressions are often not the best way to deal with linguistic data. Perl is also a bit slow for heavy parsing and segmenting -- especially if you use Parse::RecDescent ;) -- but it's definitely a good place to start.

@INBOOK{mikheev2002text,
  chapter = {10},
  pages = {201-218},
  title = {Text Segmentation},
  publisher = {Oxford University Press},
  year = {2002},
  editor = {Ruslan Mitkov},
  author = {Andrei Mikheev},
  address = {Oxford},
}

@BOOK{mitkov2002handbook,
  title = {Handbook of Computational Linguistics},
  publisher = {Oxford University Press},
  year = {2002},
  editor = {Ruslan Mitkov},
}
[download]

--
Damon Allen Davison
http://www.allolex.net

In reply to Re: NLP - natural language regex-collections? by allolex
in thread NLP - natural language regex-collections? by erix

Are you posting in the right place? Check out Where do I post X? to know for sure.
Posts may use any of the Perl Monks Approved HTML tags. Currently these include the following:
<code> <a> <b> <big> <blockquote> <br /> <dd> <dl> <dt> <em> <font> <h1> <h2> <h3> <h4> <h5> <h6> <hr /> <i> <li> <nbsp> <ol> <p> <small> <strike> <strong> <sub> <sup> <table> <td> <th> <tr> <tt> <u> <ul>
Snippets of code should be wrapped in <code> tags not <pre> tags. In fact, <pre> tags should generally be avoided. If they must be used, extreme care should be taken to ensure that their contents do not have long lines (<70 chars), in order to prevent horizontal scrolling (and possible janitor intervention).
Want more info? How to link or How to display code and escape characters are good places to start.


P is for Practical
	PerlMonks