2010 · International Symposium on Information Management in a Changing World (IMCW 2010), Communications in Computer and Information Science 96
New Approach for Automated Categorizing and Finding Similarities in Online Persian News
Evidence basis: full-text-reviewed · Review status: catalog-reviewed; paper-author approval pending
machine-learning benchmark-datasets
Persian news text categorization document similarity tf-idf SVM Reuters web crawler PHP keyword extraction semantic similarity
Core contribution: The paper combines automated Persian-news categorization with a web system for retrieving similar news items.
Problem and motivation
Online Persian news requires automatic category assignment and retrieval of related items, but the paper notes that no standard similarity test bench or universally accepted similarity measure was available (p. 61, abstract; pp. 62-63).
Method and contribution
A crawler collects news; preprocessing removes general words and creates tf*idf-style vectors; an S-V-M categorizer trained with Reuters material assigns general categories. Similarity retrieval uses stored headline/summary/topic/date features and prioritized permutations from all features down to one keyword (pp. 63-65).
Findings and evidence
The Linux/PHP crawler and similarity finder were manually assessed on 100 news pieces and corresponding results. The authors report 79% precision when permutations of headline, summary, and topic are used (p. 66). No standard benchmark split or independent reproduction was verified.
Limitations and future directions
Limitations: Surface subject/keyword resemblance can miss semantic similarity or match documents that share words without sharing meaning. The corpus protocol, train/test split, and larger benchmark size are unknown; the public copy does not establish them.
Future work: Add semantic features representing both keywords and concepts so that conceptually related Persian news can be retrieved even when surface words differ (pp. 66-67).
Sources and identifiers
- Published version published
- Public conference copy · PDF public_full_text
When to cite this paper
Cite this paper when your work uses or compares persian-news feature extraction using grammar-aware general-word removal and keyword/topic/date features.
- Persian-news feature extraction using grammar-aware general-word removal and keyword/topic/date features.
- The crawler plus tf*idf/S-V-M architecture trained with Reuters categories for automated news classification.
- Prioritized feature-permutation retrieval for related news, including the authors' 79% manual-precision result and its non-standard benchmark boundary.
- The explicit motivation for semantic similarity using concepts in addition to keywords.
Citation
@inproceedings{ezzatiJivan2010newapproach,
author = {Naser Ezzati Jivan and Mahlagha Fazeli and Khadije Sadat Yousefi},
title = {New Approach for Automated Categorizing and Finding Similarities in Online Persian News},
year = {2010},
booktitle = {International Symposium on Information Management in a Changing World (IMCW 2010), Communications in Computer and Information Science 96},
pages = {120-128},
publisher = {Springer Berlin Heidelberg},
issn = {1865-0929, 1865-0937},
isbn = {9783642160318, 9783642160325},
doi = {10.1007/978-3-642-16032-5_11},
url = {https://doi.org/10.1007/978-3-642-16032-5_11}
}Other citation formats for Word and reference managers
Jivan, N. E., Fazeli, M., & Yousefi, K. S. (2010). New Approach for Automated Categorizing and Finding Similarities in Online Persian News. In International Symposium on Information Management in a Changing World (IMCW 2010), Communications in Computer and Information Science 96 (pp. 120-128). https://doi.org/10.1007/978-3-642-16032-5_11N. E. Jivan, M. Fazeli, and K. S. Yousefi, "New Approach for Automated Categorizing and Finding Similarities in Online Persian News," in International Symposium on Information Management in a Changing World (IMCW 2010), Communications in Computer and Information Science 96, pp. 120-128, 2010, doi: 10.1007/978-3-642-16032-5_11