PROJECT
Taiwanese Cultural Benchmark
A benchmark-oriented project for evaluating culturally grounded language understanding with a focus on Taiwanese contexts and knowledge.

Principal Investigator
Prof. Shu-Kai Hsieh is a joint-appointed professor at the Graduate Institute of Linguistics and the Institute of Brain and Mind Sciences at National Taiwan University. His work sits at the intersection of computational linguistics, language resources, and the quest to understand how language shapes — and is shaped by — cognition, culture, and consciousness. He leads the LOPE Laboratory and is the principal developer of the Chinese Wordnet (CWN). Outside research, he enjoys reading Buddhist scriptures, playing instruments, and practising Chinese calligraphy.
PROJECT
A benchmark-oriented project for evaluating culturally grounded language understanding with a focus on Taiwanese contexts and knowledge.
PROJECT_CWN
Chinese Wordnet (CWN) is a long-running lexical knowledge resource for representing Chinese word senses and lexical-semantic relations through synsets, glosses, examples, and a connected semantic network.

MultiMoCo
A pioneering large-scale multimodal corpus for languages in Taiwan that integrates video, dialogue, caption, and gesture layers with human annotation and multimodal machine learning workflows.
Natural Language Processing (First View) 2026
This study investigates the nuanced challenges of fine-grained word sense disambiguation (WSD) tasks with regular polysemy detection (RPD) of the named entity, focusing on evaluating the trade-offs between encoder and decoder-based model performance and computational efficiency. The datasets, including Chinese Wordnet 2.0 (CWN) as sense inventory, the Social Media Corpus (PTT) for user-generated content, and the Academia Sinica Balanced Corpus (ASBC) for formal linguistic data, were chosen to provide a diverse and representative framework for evaluating both common nouns and proper nouns with regular polysemy in Taiwan Mandarin. This analysis evaluated ten encoder- and decoder-based models, assessing their performance on two tasks. The encoder-based models demonstrate comparable accuracy to the decoder-based models on WSD tasks (77.5% vs. 78.5%), and similarly strong performance in RPD tasks (84.2% vs. 83.8%). On a large-scale all-words WSD task, the encoder model not only outperformed the decoder model but also generated substantially lower carbon emissions – an eight-fold reduction. These differences underscore the trade-offs between model architecture and task-specific performance, highlighting the necessity for balancing performance and energy efficiency in the design and application of language models, advocating for sustainable and eco-friendly practices in natural language processing development.
Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), San Diego, California 2026
This paper investigates how large language models internally process nonliteral language. Analyzing five categories spanning slang, metaphor, and idioms across all 48 layers of Gemma-3-12B-IT with Gemma Scope 2 sparse autoencoders, we find a lexical familiarity gradient: processing depth depends on available prior lexical knowledge, not figurative type. Idioms diverge at L1 as entrenched units; expressions built from familiar words (metaphors, semantic-shift and constructional slang) converge at L7–9; neologisms peak at L41, activating 3× more unique features. Paraphrase residual analysis confirms strong signals only at the gradient endpoints, yielding a three-tier hierarchy of entrenched retrieval, known-word reanalysis, and novel-word construction. Crucially, this peak-layer structure replicates in base models (Gemma-PT, Qwen-Base), demonstrating that the gradient is a robust property of pretrained representations rather than an alignment artifact. We additionally identify an activation density confound in SAE feature counts that produces spurious cross-condition convergence. Overall, processing depth is better predicted by lexical familiarity than by figurative type, with implications for robustness to non-standard language and for SAE-based interpretability.
Intelligent Computing (Computing Conference 2026) 2026
This study introduces CorPilot, an innovative multi-agent framework that integrates Large Language Models (LLMs) to function as a form of collaborative AI, designed to automate and streamline corpus linguistic research. By structuring interactions between specialized agents within this agentic framework, CorPilot assists with tasks such as querying, annotation, semantic classification, and data analysis. Demonstrated through an empirical case study of Mandarin Chinese constructions, CorPilot replicates existing manual annotation results and reveals finer semantic distinctions and novel linguistic patterns previously unattainable through traditional methods. Our findings illustrate that CorPilot represents a significant methodological advancement, addressing challenges in scalability, replicability, and interpretability of corpus linguistics research. The framework's modular design also facilitates future extensions into various linguistic domains, holding considerable potential for theoretical and practical advancements in linguistics and computational humanities.
Language Resources and Evaluation Conference (LREC) 2026 2026
Hyperbolic embeddings such as the Poincaré model effectively represent lexical hierarchies with low distortion, yet their cross-lingual generalizability remains largely unexplored. This study investigates cross-lingual transfer by training 20-dimensional Poincaré embeddings exclusively on Open English WordNet (OEWN) hypernymy relations and evaluating on aligned Chinese Wordnet (CWN) synsets under a vocabulary-constrained transfer setting, where CWN-relevant synsets appear in OEWN training data but no Chinese-language supervision is used. We report robust statistical evidence based on the final 10 training checkpoints: Poincaré embeddings achieve 2.57× higher Mean Reciprocal Rank (MRR) than Euclidean embeddings on CWN (0.030 ± 0.001 vs 0.012 ± 0.000, p < 0.001, Cohen’s d = 34.48) and 5.61× higher on OEWN (0.016 ± 0.000 vs 0.003 ± 0.000, p < 0.001, d = 42.48). Furthermore, hierarchical filtering leveraging the radial dimension of hyperbolic space provides substantial additional gains: +74.6% MRR improvement on CWN and +25.8% on OEWN (both p < 0.001). The model achieves higher absolute performance on the zero-shot CWN test set (MRR = 0.052 ± 0.002) than on the in-domain OEWN test set (MRR = 0.020 ± 0.001). We attribute this to structural alignment: CWN’s broader branching factor (4.32 vs 1.10) and moderate depth naturally suit hyperbolic geometry’s capacity to compactly represent hierarchies. Our findings demonstrate that geometric properties learned from English hypernymy transfer robustly across languages when semantic structures align. We release the aligned CWN–OEWN hypernymy evaluation dataset and complete evaluation framework to facilitate future research on geometry-based cross-lingual semantic modeling.
Reference Module in Social Sciences 2026
Multimodal lexical items (MLIs) are conceptual entities whose meanings emerge from the dynamic integration of multiple communicative channels (e.g., text, gesture, imagery, prosody). This chapter provides an interdisciplinary, data-driven overview of how researchers from corpus linguistics, computational linguistics, neuro-psycholinguistics, and cognitive science collectively explore MLIs. We survey frameworks related to MLI semantic representation (e.g. Multimodal Distributional Semantics and Multimodal Semantics for Affordances and Actions), multimodal context of MLIs (e.g., Multimodal Construction Grammar), and multimodal lexical resources (e.g., Frame2). By examining theoretical constructs, like embodiment and affordances, and empirical methodologies, such as corpus annotation, machine learning, and experimental research, we underscore the multifaceted nature of lexical meaning in a richly multimodal world. Ultimately, we propose future directions for a more holistic, context-sensitive approach that unites traditional linguistic structures with perceptual, embodied, and dynamic environmental cues to investigate MLIs.
Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) 2026
Journal of Library and Information Studies 2025
Large language models (LLMs) have in recent years spurred research across various sectors, owing to their remarkable zero-shot or few-shot performance. This capability has become indispensable for individuals seeking to integrate these language models into their workflows effectively. In this paper, based on in-depth linguistic analyses, we explore the application of an LLM, specifically GPT-4, in generating Chinese language textbooks tailored for grade school students. This encompasses the creation of main lesson texts alongside accompanying Chinese character exercises. Experimental results suggest that the LLM-generated textbook lessons are a viable research direction. The initial outcomes demonstrate the ability of LLM to generate texts of satisfactory quality appropriate for a specified grade level. The contributions of this work include pioneering the quantitative analysis of Chinese language textbooks for native speakers in Taiwan and leveraging an LLM to automatically generate textbook content and accompanying Chinese character exercises targeted at native Chinese speakers, which is a novel approach facilitated by the development of prompts tailored to different language learning levels. The study also conducts quantitative and qualitative comparisons between machine-generated lessons and those developed by educational professionals in Taiwan.
arXiv preprint arXiv:2504.13603 2025
The recent advances in Legal Large Language Models (LLMs) have transformed the landscape of legal research and practice by automating tasks, enhancing research precision, and supporting complex decision-making processes. However, effectively adapting LLMs to the legal domain remains challenging due to the complexity of legal reasoning, the need for precise interpretation of specialized language, and the potential for hallucinations. This paper examines the efficacy of Domain-Adaptive Continual Pre-Training (DACP) in improving the legal reasoning capabilities of LLMs. Through a series of experiments on legal reasoning tasks within the Taiwanese legal framework, we demonstrate that while DACP enhances domain-specific knowledge, it does not uniformly improve performance across all legal tasks. We discuss the trade-offs involved in DACP, particularly its impact on model generalization and performance in prompt-based tasks, and propose directions for future research to optimize domain adaptation strategies in legal AI.
Proceedings of the Workshop: Bridging Neurons and Symbols for Natural Language Processing and Knowledge Graphs Reasoning (NeusymBridge)@ LREC-COLING-2024 2024
Compressibility is closely related to the predictability of the texts from the information theory viewpoint. As large language models (LLMs) are trained to maximize the conditional probabilities of upcoming words, they may capture the subtlety and nuances of the semantic constraints underlying the texts, and texts aligning with the encoded semantic constraints are more compressible than those that do not. This paper systematically tests whether and how LLMs can act as compressors of semantic pairs. Using semantic relations from English and Chinese Wordnet, we empirically demonstrate that texts with correct semantic pairings are more compressible than incorrect ones, measured by the proposed compression advantages index. We also show that, with the Pythia model suite and a fine-tuned model on Chinese Wordnet, compression capacities are modulated by the model’s seen data. These findings are consistent with the view that LLMs encode the semantic knowledge as underlying constraints learned from texts and can act as compressors of semantic information or potentially other structured knowledge.
Frontiers in Language Sciences 2024
Formosan languages, spoken by the indigenous peoples of Taiwan, have unique roles in the reconstruction of Proto-Austronesian Languages. This paper presents a real-world Formosan language speech dataset, including 144 h of news footage for 16 Formosan languages, and uses self-supervised models to obtain and analyze their speech representations. Among the news footage, 13 h of the validated speech data of Formosan languages are selected, and a language classifier, based on XLSR-53, is trained to classify the 16 Formosan languages with an accuracy of 86%. We extracted and analyzed the speech vector representations learned from the model and compared them with 152 manually coded linguistic typological features. The comparison shows that the speech vectors reflect Formosan languages' phonological and morphological aspects. Furthermore, the speech vectors and linguistic features are used to construct a linguistic phylogeny, and the resulting genealogical grouping corresponds with previous literature. These results suggest that we can investigate the current real-world language usages through the speech model, and the dataset opens a window to look into the Formosan languages in vivo.
Proceedings of the 35th Conference on Computational Linguistics and Speech Processing (ROCLING 2023) 2023
In this research, we comprehensively analyze the potential biases inherent in Large Language Model, utilizing meticulously curated input data to ascertain the extent to which such data sway machine-generated responses to yield prejudiced outcomes. Notwithstanding recent strides in mitigating bias in LLM-based NLP, our findings underscore the continued susceptibility of these models to data-driven bias. We have integrated the PTT NTU board as our primary data source for this investigation. Moreover, our study elucidates that, in certain contexts, machines may manifest biases without supplementary prompts. However, they can be guided toward rendering impartial responses when provided with enhanced contextual nuances.
2020 International Conference on Technologies and Applications of Artificial Intelligence (TAAI) 2020
The modern conversational agent requires high-quality datasets, which are often the bottlenecks when building models. This paper introduces MatDC, an entirely human-produced dialogue dataset with full semantic annotations in Chinese. The dataset features linguistic variations given users' intents and fully annotated semantic slots. MatDC dataset was completely human-edited, and the curation comprises two stages. At first, templates design stage, domain editors first construct schemas and compose ten dialogues between the agents and the users based on the back-end database. Secondly, in the dialogue rewrite stage, rewriters generate sentential variations for each template, under the constraints that the normalized slot values are kept unchanged. The underlying methodology of the MatDC is more open to extension and more adaptable to different domains. To demonstrate the applicability of the dataset, we build a dialogue agent with conventional pipeline architecture. We expect the MatDC dataset to provide additional training data and testing ground for dialogue agent studies.
Corpus Linguistics and Linguistic Theory 2016
While much attention has been paid to the complement coercion operation in English (e.g., began a book), the same phenomenon in Chinese is still under-researched. Our study examines twenty coercing verbs in Chinese, creating a coercion profile for each verb and conducting a cluster analysis based on the coercion profiles. The results suggest that semantically related verbs in Chinese tend to have similar coercion profiles. We also identify a diverse range of nouns that can be coerced in Chinese. Finally, it is demonstrated that generative approaches to the complement coercion operation in Chinese can be complemented by cognitive-functional approaches.
International Journal of Computational Linguistics & Chinese Language Processing, Volume 18, Number 2, June 2013-Special Issue on Chinese Lexical Resources: Theories and Applications 2013
There has been no consensus as to what constitutes a set of base concepts in the mental landscape. With the aim of exploring base concepts in Chinese, this paper proposes that frequently-occurring words in the glosses of a lexical resource such as the Chinese Wordnet can be seen as a candidate set of base concepts because the glosses use basic words. The present study identified 130 base concepts in Chinese. The Base Concepts in EuroWordNet were adopted as a reference for comparison. While only 44.6% of the base concepts identified in the present study have an equivalent in the set of Base Concepts of EuroWordNet, the other base concepts extracted by our gloss-based approach also reflect a certain degree of basicness. It is hoped that both the overlap and the difference between different sets of base concepts identified in different languages and by different approaches can deepen our understanding of the basic core in the mind. Additionally, it is also hoped that the set of base concepts identified in the present study can have computational as well as pedagogical applications in the future.
// FRONTIER_RESEARCH