Skip to content

Latest commit

 

History

History

README.md

StructureDataset GramAdapt Crosslinguistic Social Contact Dataset

CLDF Metadata: StructureDataset-metadata.json

Sources: sources.bib

The GramAdapt Crosslinguistic Social Contact Dataset is a pioneering dataset with social, cultural, and demographic data about contact scenarios from across the globe. The data mostly concern interactions between a Focus language community and it's Neighbour language community, with 34 contact pairs represented. The data are qualitative with quantitative potential, where language community experts have provided best-assessment answers to questions about social contact in their communities of expertise.

property value
dc:bibliographicCitation Eri Kashima, Francesca Di Garbo, Oona Raatikainen, Rosnátaly Avelino, Sacha Beck, Anna Berge, Ana Blanco, Ross Bowden, Nicolás Brid, Joseph M Brincat, María Belén Carpio, Alexander Cobbinah, Paola Cúneo, Anne-Maria Fehn, Saloumeh Gholami, Arun Ghosh, Hannah Gibson, Elizabeth Hall, Katja Hannß, Hannah Haynie, Jerry Jacka, Matias Jenny, Richard Kowalik, Sonal Kulkarni-Joshi, Maarten Mous, Marcela Mendoza, Cristina Messineo, Francesca Moro, Hank Nater, Michelle A Ocasio, Bruno Olsson, Ana María Ospina Bozzi, Agustina Paredes, Admire Phiri, Nicolas Quint, Erika Sandman, Dineke Schokkin, Ruth Singer, Ellen Smith-Dennis, Lameen Souag, Yunus Sulistyono, Yvonne Treis, Matthias Urban, Jill Vaughan, Deginet Wotango Doyiso, Georg Ziegelmeyer, Veronika Zikmundová. (2023). GramAdapt Crosslinguistic Social Contact Dataset. (1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7508054
dc:conformsTo CLDF StructureDataset
dc:license https://creativecommons.org/licenses/by/4.0/
dcat:accessURL https://github.com/cldf-datasets/gramadapt
prov:wasDerivedFrom
  1. cldf-datasets/gramadapt v1.2
  2. Glottolog v5.3
prov:wasGeneratedBy
  1. python: 3.12.3
  2. python-packages: requirements.txt
rdf:ID gramadapt
rdf:type http://www.w3.org/ns/dcat#Distribution

The ValueTable lists answers to the questions in the GramAdapt questionnaire coded as values for the Focus group language (except for values for the two 'synthetic' parameters 'Focus language' and 'Contact pair', which may also be coded for the language of the Neighbour group).

property value
dc:conformsTo CLDF ValueTable
dc:extent 13497

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
Primary key
Language_ID string References languages.csv::ID
Parameter_ID string References parameters.csv::ID
Value string
Code_ID string References codes.csv::ID
Comment string
Source list of string (separated by ;) References sources.bib::BibTeX-key
Contactset_ID string Link to the corresponding contact set.
References contributions.csv::ID
Respondent string
property value
dc:extent 49

Columns

Name/Property Datatype Description
ID string Primary key
Name string
Editor boolean
Valid choices:
Yes No

Table media.csv

property value
dc:conformsTo CLDF MediaTable
dc:extent 90

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
Primary key
Name string
Description string
Media_Type string
Regex: [^/]+/.+
Download_URL anyURI
Path_In_Zip string

The GramAdapt dataset is constructed around contact sets, pairs of Focus andNeighbour language communities. This table lists the languages spoken by either of these communities.

property value
dc:conformsTo CLDF LanguageTable
dc:extent 68

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
Primary key
Name string
Macroarea string
Latitude decimal
≥ -90
≤ 90
Longitude decimal
≥ -180
≤ 180
Glottocode string
Regex: [a-z0-9]{4}[1-9][0-9]{3}
ISO639P3code string
Regex: [a-z]{3}

The GramAdapt dataset provides two types of contributions: (1) 'contact sets', i.e. descriptions of a contact situation between two neighbouring communities speaking different languages. Each contact set is unique in terms of the timeframe they respond for. We urge researchers who use this dataset to read the Comments column for questions CID P1, P2, and P3 carefully for each set, to get a sense of the heterogeneity of timeframes represented in each set, as well as the whole dataset. (2) 'rationales' explaning the goals, definitions and theoretical support for (sets of) questions in the GramAdapt questionnaire.

property value
dc:conformsTo CLDF ContributionTable
dc:extent 125

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
(1) For contact sets: The unique identifier of a contact pair. The two digit IDs were assigned based on order of completion. Sets that are linked by language communities, but represent different time slices, contain a Roman alphabet symbol, i.e. Set06a and Set06b are for the contact scenario between Maltese and Sicilian, but for different time periods of contact. (2) For rationales the identifier are based on domain and question number.
Primary key
Name string
Description string The rationale document formatted as CLDF Markdown.
Contributor string
Citation string
Type string
Valid choices:
contactset rationale
Focus_Language_ID string Link to the language spoken by the focus group in a contact set.
References languages.csv::ID
Neighbour_Language_ID string Link to the language spoken by the neighbour group in a contact set.
References languages.csv::ID
Author_IDs list of string (separated by ) References contributors.csv::ID
Reviewer_IDs list of string (separated by ) The GramAdapt team members responsible for checking the responses.
References contributors.csv::ID
Area string Information about the geographical location of each contact set based on Autotyp areal classification
Document string Link to rendered Markdown document.
References media.csv::ID
Source list of string (separated by ;) References sources.bib::BibTeX-key

This table lists the questions of the GramAdapt questionnaire with links to the respective rationale.

property value
dc:extent 265

Columns

Name/Property Datatype Description
ID string Primary key
Name string
Rationale list of string (separated by ) Link to the rationale for this question.
References contributions.csv::ID

Questions in GramAdapt may be broken up into several sub-questions. This table lists the atomic sub-questions.

property value
dc:conformsTo CLDF ParameterTable
dc:extent 392

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
Primary key
Name string
Description string
ColumnSpec json
datatype string
Regex: Binary-YesNo|Comment|Scalar|Types|TypesSequential|TypesMultiple|Value
Binary-YesNo: A binary answer of either ‘Yes’ or ‘No; Comment: Not preset, just a comment field, i.e. free response. Scalar: A Likert 5 point scale. The response is in textual form, but represents ordinals on a scale of 1-5 (e.g. “Neither positive nor negative” -> 3). Types: A list of preset answers where only one can be chosen (e.g. “FL, NL, Some other language, This is highly contextual”) TypesSequential: A list of ordered, categorical answers where only one can be chosen Types-Multiple: A list of preset answers, where multiple can be chosen Value: A numerical value
Question_ID string Links to the corresponding GramAdapt question.
References questions.csv::ID
Domain string
Valid choices:
OV DEM DFK DKN DLB DLC DTR
Indicating the domain to which the responses apply. Possible options are the overview questionnaire (OV) and social domains (DEM = Exchange and Marriage; DFK = Family and Kin; DKN = Knowledge; DLB = Labour; DLC = Local Community; DTR = Trade).
Tag string
Valid choices:
P D S B O I T E OD OG OI OL OS OH OE OC OB OT
Questions of the overview questionnaire may be tagged as OD = Demographics; OG = Language geography; OI = Language and identity; OL = Literacy; OS = Social structure; OH = History; OE = Respondent fieldwork experience; OC = Response confidence; OB = Behaviour affecting biases; OT = Time frame. Questions of the domains questionnaire may be tagged as P = Preamble; D = Domain characterisation; S = Social network; B = Behaviour affecting biases; O = Linguistic output of Focus group people; I = Linguistic input of Focus group people, i.e. the output of Neighbour group people; T = Language transmission to children; E = Ending questions about data source and confidence.
Is_Timeframe_Comment boolean
Valid choices:
yes no
Flag signaling whether answers to this question describe a timeframe.
Use_Equivalence boolean
Valid choices:
yes no
Questions pertaining to whether language use is equivalent or language contact dynamics between Focus and Neighbour Groups are used in equivalent ways or not. The tag relates to questions of language uses in social domains, and discrepancies in fluency between Focus and Neighbour group people when speaking.
Socio-Political_Power boolean
Valid choices:
yes no
Questions on whether there are differences in socio-political power between the Focus and Neighbour groups.
Language_Loyalty boolean
Valid choices:
yes no
For the purpose of this dataset language loyalty is defined as a tendency to be loyal to one’s language, typically by expressing a desire to retain an identity that is expressed through the use of that language. This multicausal factor thus significantly overlaps with "Use Equivalence".
Attitudes_and_Ideologies boolean
Valid choices:
yes no
Concerns the two groups’ attitudes towards each other in general.

Table codes.csv

property value
dc:conformsTo CLDF CodeTable
dc:extent 1524

Columns

Name/Property Datatype Description
ID string
Regex: [a-zA-Z0-9_\-]+
Primary key
Parameter_ID string The parameter or variable the code belongs to.
References parameters.csv::ID
Name string
Description string
Ordinal integer
≥ 1
≤ 5
Ordinal, representing the value for a Scalar parameter on a 5 point Likert scale.