CÁC BÀI BÁO KHOA HỌC 23:50:24 Ngày 25/04/2024 GMT+7
An experimental study on lexicalized statistical parsing for Vietnamese

Syntactic parsing is a central problem and a challenge in the field of natural language processing. It attracts many studies and consequently there exists the effective parsers for several popular languages such as English and Chinese. For Vietnamese parsing, there have been a few studies focusing on this problem, these studies lack of applying modern techniques, and no popular parser has been released. This paper presents the first study on developing a Vietnamese wide coverage parser based on lexicalized probabilistic context free grammar (LPCFG) and using a standard parsed corpus (similar to Penn Treebank). In this paper the Bikel's parser is modified to analyze Vietnamese. We also provide a comparison based on investigating different parsing models and different linguistic features. The best configuration achieves around 78% of F-score. © 2009 IEEE.


 Le A.-C., Nguyen P.-T., Vuong H.-T., Pham M.-T., Ho T.-B.
   275.pdf    Gửi cho bạn bè
  Từ khóa : Central problems; Experimental studies; F-score; Linguistic features; NAtural language processing; Parsed corpora; Probabilistic context free grammars; Syntactic parsing; Treebanks; Computational linguistics; Context free grammars; Knowledge engineering; Natural language processing systems; Query languages; Standardization; Systems engineering; Formal languages