logo CoMeRe

This page: https://hdl.handle.net/11403/comere/cmr-88milsms
Back to Repository main page: https://hdl.handle.net/11403/comere

Corpus CoMeRe cmr-88milsms-tei-v1 :
88milSMS. A corpus of authentic text messages in French.

logo Ortolang
Open Resources and TOols for LANGuage

How to cite this resource

Panckhurst R., Détrie C., Lopez C., Moïse C., Roche M., Verine B. (2016). 88milSMS. A corpus of authentic text messages in French (nouvelle version du corpus ISLRN : 024-713-187-947-8). In Chanier T. (ed) Banque de corpus CoMeRe. Ortolang : Nancy. [cmr-88milsms-tei-v1 ; https://hdl.handle.net/11403/comere/cmr-88milsms/cmr-88milsms-tei-v1]


The first version of the corpus (ISLRN : 024-713-187-947-8) was produced in 2014 as part of the "sud4science LR project". More than 88,000 authentic SMS, sent by hundreds of donators living mainly in the Montpellier area, were collected, in 2011, then anonymised, by the researchers, their student interns and a legal adviser-CIL.

The initial corpus was then converted to TEI standard in the project CoMeRe (Communication Médiée par les Réseaux). This project aims to build a kernel corpus assembling existing corpora of different CMC (Computer-Mediated Communication) genres and new corpora build on data extracted from the Internet. These heterogenous corpora will be structured and processed in a uniform way, complemented with metadata. CoMeRe will be released as OpenData through the national infrastructure Ortolang, following constraints which will be reused for the forthcoming “Corpus de Référence du Français”. Project supported by the national consortium Corpus-écrits, sub-part of Huma-Num, and Ortolang (French correspondant to DARIAH)

Keywords: Short Message Service; Computer Mediated Communication; CMC;


This corpus contains :



This corpus can be freely distributed and shared subject only to attribution. The way to reference / cite the corpus is given in the bibliographicCitation