arXiv
Open-access scholarly preprint repository for global research communities.
Available information varies by company and source.
Profile record updated:
Company facts
- Entity type
- COMPANY
- Founded
- 1991
- Headquarters
- United States
- Company size
- 10–49
- Market role
- Publisher & Media Owner
- Official website
- arxiv.org
What arXiv does
arXiv operates nonprofit publishing infrastructure for scholarly communication. It aggregates preprint submissions, applies curation and moderation processes, hosts the content for public access, and distributes it at internet scale. Value is created through trusted dissemination, long-term archive utility, domain reputation and deep integration into academic research workflows. Revenue support comes from institutional and philanthropic funding rather than reader subscriptions or advertising.
Category differentiation
arXiv is not a commercial journal publisher, adtech platform or general-purpose SaaS vendor. It is an open-access scholarly preprint repository and research distribution service.
Strategic context
AI-supported assessment from the existing company research; distinguish interpretation from sourced facts.
arXiv is an open-access scholarly preprint repository focused on physics, mathematics, computer science and adjacent research fields. It curates, hosts and distributes research manuscripts prior to formal journal publication, serving researchers, academic institutions and the broader scientific community. Following Cornell’s June 2026 announcement, arXiv became an independent nonprofit on 1 July 2026 while keeping its headquarters at Cornell Tech’s Tata Innovation Center in the United States. The organisation creates value by providing a trusted, high-volume distribution and discovery layer for research outputs. It does not operate as a conventional commercial software vendor or ad-supported media property. Its funding model is built around institutional support, philanthropic gifts and nonprofit backing, with recent capital explicitly allocated to completing migration to cloud infrastructure and supporting ongoing platform operations.
Company news briefing
Briefing updated:
arXiv continues to serve as the primary repository for foundational AI research, hosting recent breakthroughs such as OpenAI’s mathematical proofs and Microsoft’s SkillOpt framework for self-evolving agents. Analyses disseminated via the platform are increasingly influencing the dialogue on AI search transparency, behavioural fingerprinting, and the formalisation of autonomous agent security protocols. These contributions reinforce arXiv’s position as a critical infrastructure for the technical validation and standardisation of generative AI and agentic systems within the global technology sector.
Business model & monetisation
arXiv monetises through nonprofit funding mechanisms rather than transactional media sales. Its commercial support base consists of institutional membership or support contributions, foundation grants and philanthropic gifts. Access to the repository is open, so monetisation is decoupled from readership and tied instead to ecosystem support for essential research infrastructure.
- Institutional support and membership funding
- Service Fee
- Foundation grants
- Service Fee
- Philanthropic gifts and donations
- Service Fee
Products & capabilities
No products with linked sources are available in this view.
Recent recorded signals
Dates refer to the source publication. Older entries are historical context, not evidence of a new event.
Models Struggle to Invert Charts into Values
Large Language Models & Chart Understanding · Recorded impact score: 2/5
The article explains that multimodal models can describe charts well but often fail to accurately recover numeric values because reading a chart requires inverting visual encodings (length, angle, position, color) into numbers. Error modes depend on the encoding (e.g., truncated y-axes, log scales, overlapping colours, legends far from marks). The author reviews benchmarks (ChartQA, PlotQA, CharXiv), recommends separating label-reading from arithmetic in evaluations, and provides practical advice: attach raw data instead of images, request extracted values before calculations, increase image resolution, and explicitly state axis properties. An example Python snippet shows how to generate grounded evaluation charts from known data.
- Multimodal models typically describe chart content well but struggle to invert visual encodings to recover precise numeric values.
- Chart reading failure modes vary by encoding: labelled bars are easiest, unlabelled/truncated axes and log scales cause large errors, and colour/legend binding and stacked/dual-axis charts commonly produce misattribution or arithmetic mistakes.
When the Student Talked Back: LLM Distillation Shift
Large Language Models (LLM) & AI · Recorded impact score: 3/5
This essay traces how classical model distillation assumptions (fixed input distributions, teachers producing probability vectors over closed class sets, and students trained to match those vectors) were disrupted by the rise of large language models. Over roughly five years the field moved from thinking about distillation as simple compression toward 'capability transfer' — using larger models to enable smaller models to perform complex tasks. The author frames the shift as occurring in three stages, with Stage One summarized as 'Sequences Are Not Pictures', highlighting how sequence modeling for language violated earlier assumptions rooted in image classification pipelines. The piece is published in The Sequence newsletter (Substack) on 2026-07-13.
- The 2015 distillation paper assumed a fixed input distribution, a teacher producing a probability vector over a closed set of classes, and a student trained to match that vector.
- The arrival of large language models broke the core assumptions behind classical distillation pipelines.
Explore company relationships
Questions about arXiv
What is arXiv?
arXiv is an open-access repository for scholarly preprints, focused on physics, mathematics, computer science and related research fields.
Who uses arXiv?
Researchers, academics, students, universities, libraries and scientific communities use arXiv to submit, access and distribute preprints.
How does arXiv make money?
arXiv is funded through institutional support, foundation grants and philanthropic gifts rather than subscriptions or advertising.
Sources & coverage
This profile uses public, official and technically observable information. Missing information does not prove that a product or relationship does not exist. The list below does not imply that every profile statement has been verified.
7 publicly documented primary sources and citations linked across the market graph.
Continue your research on arXiv
Explorer includes additional company details, a Watchlist for up to 25 companies and your personal Strategic Intelligence Agent. It monitors your market daily and delivers tailored briefings with clear strategic context whenever relevant news occurs.
Free, with no time limit.
