\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments

Ghosh, Debanjan; Chan, Sophia

Computer Science > Computation and Language

arXiv:2410.09314 (cs)

[Submitted on 12 Oct 2024]

Title:\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments

Authors:Debanjan Ghosh, Sophia Chan

View PDF HTML (experimental)

Abstract:We present \llinstruct: An 8B instruction-tuned model that is designed to generate content for English Language Proficiency Assessments (ELPA) and related applications. Our work involves creating a new dataset of 70K instructions and explanations in the ELPA domain and using these to fine-tune Llama-3 8B models (SFT) of different sizes (e.g., SFT-17K, SFT-50K and SFT-70K). Human evaluations are conducted over unseen instructions to compare these SFT models against SOTA models (e.g., Dolly-2, Mistral, Llama-3 base version, and GPT-3.5). The findings show although all three SFT models perform comparably, the model trained on largest instruction dataset -- SFT-70K - leads to the most valid outputs ready for assessments. However, although the SFT models perform better than larger model, e.g., GPT 3.5 on the aspect of explanations of outputs, many outputs still need human interventions to make them actual ready for real world assessments.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2410.09314 [cs.CL]
	(or arXiv:2410.09314v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2410.09314

Submission history

From: Debanjan Ghosh [view email]
[v1] Sat, 12 Oct 2024 00:47:45 UTC (7,671 KB)

Computer Science > Computation and Language

Title:\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators