<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>social media analytics &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/social-media-analytics/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 14:08:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>social media analytics &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Spot Sarcasm in Punjabi, Boosting Consumer Insight</title>
		<link>https://scienmag.com/ai-learns-to-spot-sarcasm-in-punjabi-boosting-consumer-insight/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 14:08:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for sentiment analysis]]></category>
		<category><![CDATA[AI-driven cultural context interpretation]]></category>
		<category><![CDATA[consumer insight]]></category>
		<category><![CDATA[consumer insight through AI]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for sarcasm recognition]]></category>
		<category><![CDATA[hybrid mT5-LSTM model]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[mT5]]></category>
		<category><![CDATA[multilingual sarcasm detection models]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[Punjabi]]></category>
		<category><![CDATA[Punjabi language NLP datasets]]></category>
		<category><![CDATA[Punjabi social media text analysis]]></category>
		<category><![CDATA[sarcasm detection]]></category>
		<category><![CDATA[Sarcasm detection in Punjabi language]]></category>
		<category><![CDATA[sarcasm understanding in underrepresented languages]]></category>
		<category><![CDATA[sentiment analysis]]></category>
		<category><![CDATA[sentiment analysis challenges in written text]]></category>
		<category><![CDATA[social media analytics]]></category>
		<category><![CDATA[social media sentiment mining]]></category>
		<category><![CDATA[transformer-based learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=195099</guid>

					<description><![CDATA[A hybrid transformer-LSTM model detects sarcasm in Punjabi social media text with 94.29 percent accuracy, revealing hidden customer dissatisfaction for businesses.]]></description>
										<content:encoded><![CDATA[<p>Sarcasm is one of the most slippery features of human communication, and nowhere is it harder to pin down than in plain written text. When someone writes that a delayed flight was exactly how they hoped to spend their evening, the literal words say one thing while the intended meaning says the opposite. Human readers resolve this contradiction effortlessly using tone of voice, facial expression and shared cultural context, none of which survive the journey into a typed message. For businesses that mine social media to understand their customers, this gap between what is written and what is meant can quietly corrupt entire sentiment dashboards, turning complaints into praise and frustration into apparent satisfaction.</p>
<p>A new study published in Information Systems Frontiers by Jasleen Kaur and Amarjit Gill of the Edwards School of Business at the University of Saskatchewan tackles this problem for Punjabi, a language spoken by well over a hundred million people yet strikingly underrepresented in sarcasm research. The work is significant for two reasons. First, it introduces PunjSarc, a new labelled dataset of sarcastic and non-sarcastic Punjabi text drawn from blogs, news websites and social media platforms. Second, it proposes a hybrid deep learning architecture, Hybrid mT5-LSTM, that combines a multilingual transformer encoder with a recurrent sequence model, and demonstrates that this pairing outperforms existing benchmark methods for sarcasm detection in the language.</p>
<p>The researchers assembled 5,427 instances of Punjabi text, each labelled as either sarcastic or non-sarcastic, creating what they describe as the first substantial resource of its kind for the language. Because sarcastic examples are inherently rarer than straightforward statements in most collections, the dataset suffered from class imbalance, a well-known pitfall that biases classifiers toward the majority class. To counter this, the team applied a back-translation-based data augmentation strategy, in which sentences are machine-translated into another language and back again, producing paraphrases that preserve meaning while adding linguistic variety. This process expanded the dataset to 5,775 instances and allowed the authors to rigorously test whether augmentation actually helps or hurts downstream classification.</p>
<p>The technical pipeline draws on three families of textual features. The simplest is the bag-of-words representation, in which a document is reduced to the multiset of words it contains, weighted either by raw term frequency or by term frequency-inverse document frequency, a scheme that down-weights common words and highlights distinctive ones. These sparse representations have long been the workhorses of text classification because they are fast and interpretable. The second family consists of dense word embeddings, continuous vectors that place semantically similar words near each other in a learned geometric space. The third family is the transformer encoder at the heart of the hybrid model, which processes entire sequences of tokens at once and captures long-range dependencies that older models miss.</p>
<p>mT5, the multilingual variant of the T5 text-to-text transformer, was pretrained on a vast multilingual corpus and therefore carries useful linguistic knowledge about low-resource languages such as Punjabi. On its own, however, a transformer produces contextualized token representations that still need to be distilled into a single decision. The authors addressed this by feeding the transformer&#8217;s output into a long short-term memory network, a recurrent architecture introduced in 1997 that uses gated memory cells to decide what information to retain or discard as it moves through a sequence. The LSTM layer, in effect, reads the transformer&#8217;s contextual representations sequentially, capturing directional patterns in how sarcastic cues unfold within a sentence before a final classification layer renders its verdict.</p>
<p>To contextualize the hybrid model&#8217;s performance, the study ran four families of experiments on both the original and the augmented datasets: classical baseline machine learning models such as decision trees and support vector machines, ensemble learning methods, standalone deep learning models including LSTM and bidirectional LSTM, and the proposed hybrid architecture. The results on augmentation were notably mixed. Back-translation augmentation produced marginal improvements for the deep learning sequence models, which benefited from the additional training examples, but it slightly degraded performance for decision trees and, more surprisingly, for the hybrid model itself. This finding is a valuable caution for practitioners: augmentation is not a free lunch, and its effects depend heavily on how a model represents and generalizes from text.</p>
<p>The headline result concerns the hybrid model trained on the original, unaugmented data. Hybrid mT5-LSTM achieved 94.29 percent accuracy in distinguishing sarcastic from non-sarcastic Punjabi text, surpassing all benchmark models tested in the study and setting a new state of the art for the language. Because a single accuracy figure can be misleading, the authors validated the superiority of their model using a statistical t-test, providing formal evidence that the performance gain is not an artifact of random variation across runs. For a language that until now had virtually no dedicated sarcasm detection research, the jump is a substantial leap forward.</p>
<p>The broader significance of the work extends well beyond the leaderboard. Sarcasm fundamentally distorts sentiment analysis: a sarcastic review that reads as praise on the surface often encodes deep dissatisfaction underneath. Prior research has shown that sentiment classifiers that ignore sarcasm misclassify a meaningful fraction of opinionated text, and marketing scholars have documented how media sentiment shapes firm sales growth and how social media has become central to consumer behaviour and new product development. Accurate sarcasm detection therefore functions as a quality-control layer for consumer analytics, and the authors argue that business owners who deploy it can uncover hidden customer dissatisfaction that would otherwise remain invisible, enabling tailored recommendations that strengthen satisfaction and loyalty.</p>
<p>There is also a cultural and linguistic equity dimension. Most sarcasm detection systems have been built for English, with growing bodies of work on Arabic, Chinese, Hindi, Bengali, Kannada, Telugu, Tamil and Indonesian. Punjabi, despite its enormous speaker base and vibrant social media presence, has been left largely on the sidelines, in part because building labelled datasets is expensive and requires native-speaker annotation. The release of the PunjSarc dataset, together with publicly available code, lowers the barrier for other researchers and sets the groundwork for future academic work on figurative language in Punjabi, including irony, humour and code-mixed text that blends Punjabi with English.</p>
<p>The study also contributes a nuanced empirical lesson about the interaction between data augmentation and model architecture. The finding that transformer-LSTM hybrids perform best on unmodified data while simpler recurrent models gain from augmentation suggests that the capacity of a model mediates how much it benefits from additional paraphrased examples. High-capacity pretrained encoders may already encode the variability that augmentation injects, making synthetic data redundant or even noise-inducing for such models, whereas weaker learners gain genuine signal from the expanded sample. As organizations increasingly fine-tune large multilingual models for niche languages and specialized sentiment tasks, this interaction between augmentation strategy and architecture choice is exactly the kind of practical knowledge that determines whether a deployed system performs as promised.</p>
<p>For the growing community of researchers in computational linguistics and information systems, the message of this research is twofold. Sarcasm detection in low-resource languages is both tractable and commercially valuable, and the path forward runs through careful dataset construction, honest ablation of techniques such as augmentation, and hybrid architectures that marry pretrained multilingual knowledge with sequence-aware classification. Kaur and Gill&#8217;s hybrid model, validated statistically and benchmarked against a broad set of machine learning, ensemble and deep learning competitors, offers a template that can be adapted to other underserved languages. As social media conversation shifts decisively toward regional languages, tools that can read between the lines will become essential instruments for understanding what consumers everywhere are really saying.</p>
<p><strong>Subject of Research:</strong> Sarcasm detection in Punjabi social media text using a hybrid mT5-LSTM deep learning model.</p>
<p><strong>Article Title:</strong> From Sarcasm to Consumer Insight: Leveraging mT5-LSTM for Analyzing Punjabi Social Media Behavior</p>
<p><strong>Article References:</strong> From Sarcasm to Consumer Insight: Leveraging mT5-LSTM for Analyzing Punjabi Social Media Behavior. (n.d.). <a href="https://doi.org/10.1007/s10796-026-10812-5" rel="noopener noreferrer">https://doi.org/10.1007/s10796-026-10812-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10796-026-10812-5" rel="noopener noreferrer">10.1007/s10796-026-10812-5</a></p>
<p><strong>Keywords:</strong> sarcasm detection, Punjabi, natural language processing, deep learning, mT5, LSTM, data augmentation, sentiment analysis, consumer insight, machine learning, transformer-based learning, social media analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">195099</post-id>	</item>
		<item>
		<title>AI Reveals What Employees Really Think About Pay and Benefits on LinkedIn</title>
		<link>https://scienmag.com/ai-reveals-what-employees-really-think-about-pay-and-benefits-on-linkedin/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 04:06:47 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[000 LinkedIn comments]]></category>
		<category><![CDATA[AI-driven insights into employee satisfaction]]></category>
		<category><![CDATA[analysis of 42]]></category>
		<category><![CDATA[benefits packages and employee engagement]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[compensation]]></category>
		<category><![CDATA[contextual understanding of employee benefits perceptions]]></category>
		<category><![CDATA[employee benefits]]></category>
		<category><![CDATA[Employee perceptions of pay and benefits]]></category>
		<category><![CDATA[employee sentiment analysis on LinkedIn]]></category>
		<category><![CDATA[human resource management]]></category>
		<category><![CDATA[impact of workplace recognition and development]]></category>
		<category><![CDATA[LinkedIn]]></category>
		<category><![CDATA[modern workforce valuation shifts]]></category>
		<category><![CDATA[named entity recognition]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in HR research]]></category>
		<category><![CDATA[real-time feedback on total rewards]]></category>
		<category><![CDATA[sentiment analysis]]></category>
		<category><![CDATA[social media analysis of workplace perks]]></category>
		<category><![CDATA[social media analytics]]></category>
		<category><![CDATA[social media mining for employee opinions]]></category>
		<category><![CDATA[topic modeling]]></category>
		<category><![CDATA[total rewards]]></category>
		<category><![CDATA[work-life balance]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193590</guid>

					<description><![CDATA[Researchers used BERT-based natural language processing on 42,852 LinkedIn comments to reveal that employee discourse shifted sharply toward work-life balance while sentiment on total rewards varied significantly across industries.]]></description>
										<content:encoded><![CDATA[<p>When employees talk about their pay, benefits, and workplace perks, they rarely hold back on professional social media. Now, researchers have shown that this public chatter can be systematically mined with artificial intelligence to reveal how workers truly feel about what their employers offer them. A new study published in Discover Global Society analyzed more than 42,000 LinkedIn comments posted throughout 2023, using natural language processing to decode employee perceptions of total rewards—the holistic bundle of compensation, benefits, work-life balance, recognition, and career development that organizations provide. The findings offer some of the most granular real-time evidence yet on how the modern workforce evaluates its employers, and they suggest a striking shift in what employees value most.</p>
<p>The research team, Erkut Altindağ of Doğuş University and Özge Gül of Istanbul Rumeli University, collected 42,852 public LinkedIn comments from January to December 2023. They targeted discussions on posts tagged with keywords such as compensation, benefits, salary, and total rewards, along with conversations in professional human resources groups and responses to corporate announcements about benefits packages. Rather than relying on the structured questionnaires and interviews that have long dominated compensation research—methods vulnerable to social desirability bias and flawed retrospective recall—the researchers tapped into unfiltered, spontaneous discourse. They are careful, however, to frame social media as a complement to surveys rather than a replacement, acknowledging that LinkedIn participation is shaped by self-selection, vocal minorities, and the performative nature of public posting.</p>
<p>Technically, the study deployed a multi-modal analytical pipeline built on some of the most powerful tools in modern computational linguistics. At its core was a BERT-based sentiment analysis model, fine-tuned on 1,000 manually labeled compensation-related comments to classify each remark as positive, neutral, or negative. The BERT architecture, a deep bidirectional transformer encoder, outperformed simpler baselines including a TF-IDF logistic regression model (F1 = 0.71) and a lexicon-based VADER approach (F1 = 0.65), achieving a validation accuracy of 0.89 and an F1-score of 0.87 on the test set. Class imbalance in the training data was handled through class-weighted cross-entropy loss, and the framework incorporated a custom lexicon of compensation-specific terms to sharpen accuracy in this particular domain.</p>
<p>Alongside sentiment analysis, the researchers applied two complementary techniques for uncovering structure in the discourse. Topic modeling was performed using both Latent Dirichlet Allocation, optimized to 15 topics through coherence scoring, and BERTopic, a neural approach capable of tracking how themes evolved dynamically across the year. In parallel, a custom named entity recognition model built on spaCy&#8217;s framework and trained on 2,500 annotated comments extracted specific entities from the text: organizations, benefit types, and job roles. The entity recognition model achieved an F1-score of 0.83, with particularly strong performance in identifying organizations (F1 = 0.86) and benefit types (F1 = 0.83). Data quality was protected through rigorous preprocessing—English-language filtering using langdetect, deduplication via exact and fuzzy matching based on Levenshtein distance, and thorough text cleaning—while validation included inter-rater reliability assessment by three independent coders, who achieved a Cohen&#8217;s Kappa of 0.82, and five-fold cross-validation across all models.</p>
<p>The results paint a vivid picture of sector-specific sentiment in the world of work. Technology sector discussions maintained the highest positive sentiment, with a mean score of 0.42, followed by healthcare at 0.31 and financial services at 0.25. A one-way analysis of variance confirmed that these differences between industries were statistically significant (F(2, 42,849) = 23.45, p &lt; 0.001). Sentiment also fluctuated throughout the year, with notable peaks coinciding with major corporate benefits announcements, and the technology sector showed greater volatility than other industries. The authors note that because LinkedIn comments are nested within users, posts, and organizations, the standard ANOVA violates independence assumptions, so the reported statistic should be treated as a first-order approximation, with multilevel modeling recommended for future work.</p>
<p>Perhaps the most striking finding concerns how the themes of employee conversation shifted over the course of a single year. Work-life balance discussions surged from 14.2 percent of the discourse in the first quarter of 2023 to 41.2 percent by the fourth quarter—a 27 percentage-point increase in thematic prevalence. Meanwhile, traditional base compensation discussions declined by 12.0 percent. A chi-square test confirmed the statistical significance of the overall thematic shift (χ²(4) = 156.23, p &lt; 0.001), though the researchers emphasize that the omnibus test does not establish the significance of any single theme&#8217;s change; the individual percentage shifts are presented as descriptive magnitudes pending confirmatory pairwise comparisons with multiple-testing correction. Even with those caveats, the direction is clear: employees are talking less about raw salary and more about time, flexibility, and well-being.</p>
<p>The named entity recognition analysis added another layer of insight by linking specific benefits to sentiment outcomes. The model identified 3,427 unique organizations, 892 granular benefit-related entities—which aggregate onto roughly 47 canonical benefit categories—and 1,245 job roles across the dataset. Health insurance emerged as the benefit type most strongly correlated with positive sentiment (r = 0.42, p &lt; 0.001, 95 percent confidence interval 0.38 to 0.46), followed closely by flexible work arrangements (r = 0.38, p &lt; 0.001). These correlations remained robust after controlling for industry sector and temporal variation. Health insurance dominated the discourse with 12,458 mentions, while retirement benefits, despite lower frequency at 7,892 mentions, maintained a moderate positive correlation with sentiment (r = 0.31, p &lt; 0.001).</p>
<p>The theoretical backbone of the study is Social Exchange Theory, the classic framework introduced by Peter Blau in 1964, which holds that workplace relationships operate on reciprocity: employees weigh the benefits they receive against the effort and commitment they contribute. When the exchange feels fair, workers respond with engagement, loyalty, and discretionary effort; when they feel undervalued, dissatisfaction and turnover intentions follow. Viewed through this lens, the findings suggest that employees increasingly interpret non-monetary rewards—flexible schedules, development opportunities, and work-life balance initiatives—as signals of organizational commitment to their well-being, not merely as transactional extras. The observed migration of discourse away from traditional compensation and toward holistic well-being aligns with this relational interpretation and with prior survey-based evidence that workers increasingly prioritize intangible benefits over pay alone.</p>
<p>The authors are candid about the limitations of their approach. The single-platform focus on LinkedIn may miss perspectives prevalent elsewhere; the one-year observation window means the quarter-to-quarter thematic shifts should be read as within-year movements rather than stable longitudinal trends; the English-only filter introduces cultural and linguistic bias; and LinkedIn&#8217;s user base skews toward certain professional and demographic groups. The dataset itself, consisting of publicly available comments collected under applicable data protection rules, cannot be shared in raw form, though anonymized and aggregated data are available on reasonable request. Despite these constraints, the study establishes a validated, replicable framework for social media analytics in compensation research—one the authors say achieved robust overall performance (F1 = 0.83) across the pipeline.</p>
<p>The practical implications for employers are considerable. As organizations compete for talent in a tight labor market, the study suggests that total rewards strategies built around pay alone may be increasingly out of step with workforce expectations. Compensation professionals now have evidence that flexible work arrangements and health benefits generate the strongest positive sentiment, that sentiment varies meaningfully by industry, and that the timing of benefits announcements visibly moves the needle on employee discourse. The researchers call for cross-platform validation, longitudinal studies spanning multiple years, multilingual analysis capabilities, and the integration of demographic variables to capture preference variation across employee segments. They even point toward predictive models capable of anticipating emerging compensation trends before they fully materialize in public conversation. In an era when employee voice is amplified, searchable, and machine-readable, the silent signals of the workforce are silent no more—and organizations that learn to listen computationally may gain a decisive edge in designing rewards that resonate.</p>
<p>Beyond its immediate findings, the study sits within a broader methodological turn in organizational research. Computational text analysis has gained traction across management science because it captures behavior in natural settings, sidestepping the artificiality of laboratory tasks and the recall problems of retrospective questionnaires. The choice of BERT is significant in this respect: unlike earlier bag-of-words techniques that ignore word order, transformer models process each word in relation to its surrounding context, allowing them to distinguish, for example, a sarcastic complaint about a benefits package from a sincere endorsement using nearly identical vocabulary. This contextual sensitivity matters greatly in compensation discourse, where negation, hedging, and irony are common.</p>
<p>The domain-specific adaptations the researchers made also illustrate a key lesson for applied text analytics. Off-the-shelf sentiment tools are typically trained on general web text or product reviews, where the language of workplace compensation is underrepresented. By supplementing the model with a custom lexicon of compensation terms and fine-tuning on manually labeled comments, the team addressed the vocabulary gap that often degrades performance when general-purpose models are applied to specialized professional discourse. The reported gap between the BERT model and the lexicon-based VADER baseline underscores how much accuracy can be lost without such adaptation.</p>
<p>The theoretical framing also deserves emphasis. Social Exchange Theory, as elaborated by scholars such as Gould-Williams and Davies, holds that employees interpret rewards not merely as transactional payments but as signals of how much the organization values them. The study&#8217;s finding that health insurance and flexible work arrangements carry the strongest positive sentiment fits this account: benefits that touch on security and personal autonomy may function as especially potent signals of organizational care. Conversely, the decline in base-pay discussion suggests that salary, while foundational, may be increasingly treated as a baseline expectation rather than a differentiator among employers.</p>
<p>For researchers, the study also highlights unresolved measurement questions. Because sentiment scores were aggregated across comments nested within users, posts, and organizations, the effective sample size for industry comparisons is smaller than the raw comment count suggests, and future multilevel designs could partition variance at each level. Extending the framework to multilingual corpora would be particularly valuable given that compensation norms and benefit expectations vary substantially across national labor markets. If subsequent work confirms these patterns across platforms and languages, social media analytics could become a routine complement to engagement surveys, giving organizations a near real-time barometer of how their reward strategies are actually landing with the workforce.</p>
<p><strong>Subject of Research:</strong> Natural language processing analysis of employee perceptions of total rewards through LinkedIn discourse</p>
<p><strong>Article Title:</strong> Employee perceptions of total rewards revealed through natural language processing of LinkedIn discourse</p>
<p><strong>Article References:</strong> Altindağ, E., &amp; Gül, Ö. (2026). Employee perceptions of total rewards revealed through natural language processing of LinkedIn discourse. <em>Discover Global Society, 4</em>(1), Article 238. <a href="https://doi.org/10.1007/s44282-026-00565-6" rel="noopener noreferrer">https://doi.org/10.1007/s44282-026-00565-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44282-026-00565-6" rel="noopener noreferrer">10.1007/s44282-026-00565-6</a></p>
<p><strong>Keywords:</strong> natural language processing, total rewards, LinkedIn, sentiment analysis, employee benefits, compensation, social media analytics, BERT, topic modeling, named entity recognition, work-life balance, human resource management</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193590</post-id>	</item>
	</channel>
</rss>
