# Project Hadron
## Overview
**Project Hadron** is an open-source application framework for in-memory preprocessing, where
data analysis, machine learning, and other data-intensive tasks require efficiency and speed.
With :Apache Arrow as its canonical, and a more directed use of pandas,
**Project Hadron** offers effective data management, extensive interoperability, improved memory
management and hardware optimization.
At its concept, **Project Hadron** was conceived with a desire to improve the availability of
objective relevant data, increase the transparency and traceability of data lineage and facilitate
knowledge transfer, retrieval and reuse.
At its core **Project Hadron** is a selection of capabilities that
represent an encapsulated set of actions that act upon a given set of features or dataset. An
example of this would be FeatureSelection, a capability class, encapsulating cleaning data by
removing uninformative columns.
For the complete documentation [read-the-docs](https://discovery-capability.readthedocs.io/en/latest/)
## Installation
### Python version
We recommend using the latest version of Python. Project Hadron supports Python 3.8 and newer.
### Package installation
The best way to install the component packages is directly from the
[Python Package Index](https://pip.pypa.io/en/stable/) using pip.
The component package is discovery-capability and pip installed with:
```bash
pip install discovery-capability
```
if you want to upgrade your current version then using pip install upgrade with:
```bash
pip install -U discovery-capability
```
This will also install or update dependent third party packages. The dependencies are limited to
Python, PyArrow and related Data manipulation tooling such as Pandas, Numpy, scipy, scikit-learn
and visual packages matplotlib and seaborn, and thus have a limited footprint and non-disruptive
installation in a data processing environment.
## Next Steps
For next steps [read-the-docs](https://discovery-capability.readthedocs.io/en/latest/)
## License
Distributed under the MIT License. See `LICENSE.txt` for more information or reference
[MIT](https://choosealicense.com/licenses/mit/)
## Contributing
Contributions are what make the open source community such an amazing place to learn,
inspire, and create. Any contributions you make are **greatly appreciated**.
If you have a suggestion that would make this better, please fork the repo and create a
pull request. You can also simply open an issue with the tag "enhancement".
Don't forget to give the project a star! Thanks again!
1. Fork the Project
2. Create your Feature Branch (`git checkout -b feature/AmazingFeature`)
3. Commit your Changes (`git commit -m 'Add some AmazingFeature'`)
4. Push to the Branch (`git push origin feature/AmazingFeature`)
5. Open a Pull Request
Raw data
{
"_id": null,
"home_page": "https://github.com/gigas64/discovery-capability",
"name": "discovery-capability",
"maintainer": null,
"docs_url": null,
"requires_python": ">=3.8",
"maintainer_email": null,
"keywords": "data pipeline, data preprocessing, data processing pipeline",
"author": "Gigas64",
"author_email": "gigas64@opengrass.net",
"download_url": "https://files.pythonhosted.org/packages/2e/bb/442f3976408645cc5c161c344827fc545ee95eb3d01fd17384ab05216458/discovery-capability-0.23.18.tar.gz",
"platform": null,
"description": "# Project Hadron\n## Overview\n\n**Project Hadron** is an open-source application framework for in-memory preprocessing, where\ndata analysis, machine learning, and other data-intensive tasks require efficiency and speed.\nWith :Apache Arrow as its canonical, and a more directed use of pandas,\n**Project Hadron** offers effective data management, extensive interoperability, improved memory\nmanagement and hardware optimization.\n\nAt its concept, **Project Hadron** was conceived with a desire to improve the availability of\nobjective relevant data, increase the transparency and traceability of data lineage and facilitate\nknowledge transfer, retrieval and reuse.\n\nAt its core **Project Hadron** is a selection of capabilities that\nrepresent an encapsulated set of actions that act upon a given set of features or dataset. An\nexample of this would be FeatureSelection, a capability class, encapsulating cleaning data by\nremoving uninformative columns.\n\nFor the complete documentation [read-the-docs](https://discovery-capability.readthedocs.io/en/latest/)\n\n## Installation\n\n### Python version\nWe recommend using the latest version of Python. Project Hadron supports Python 3.8 and newer.\n\n### Package installation\nThe best way to install the component packages is directly from the \n[Python Package Index](https://pip.pypa.io/en/stable/) using pip.\n\nThe component package is discovery-capability and pip installed with:\n\n```bash\npip install discovery-capability\n```\n\n\nif you want to upgrade your current version then using pip install upgrade with:\n\n```bash\npip install -U discovery-capability\n```\n\nThis will also install or update dependent third party packages. The dependencies are limited to\nPython, PyArrow and related Data manipulation tooling such as Pandas, Numpy, scipy, scikit-learn\nand visual packages matplotlib and seaborn, and thus have a limited footprint and non-disruptive\ninstallation in a data processing environment.\n\n## Next Steps\nFor next steps [read-the-docs](https://discovery-capability.readthedocs.io/en/latest/)\n\n## License\nDistributed under the MIT License. See `LICENSE.txt` for more information or reference\n[MIT](https://choosealicense.com/licenses/mit/)\n\n## Contributing\n\nContributions are what make the open source community such an amazing place to learn, \ninspire, and create. Any contributions you make are **greatly appreciated**.\n\nIf you have a suggestion that would make this better, please fork the repo and create a \npull request. You can also simply open an issue with the tag \"enhancement\".\nDon't forget to give the project a star! Thanks again!\n\n1. Fork the Project\n2. Create your Feature Branch (`git checkout -b feature/AmazingFeature`)\n3. Commit your Changes (`git commit -m 'Add some AmazingFeature'`)\n4. Push to the Branch (`git push origin feature/AmazingFeature`)\n5. Open a Pull Request\n",
"bugtrack_url": null,
"license": "BSD",
"summary": "Data Science to production accelerator",
"version": "0.23.18",
"project_urls": {
"Homepage": "https://github.com/gigas64/discovery-capability"
},
"split_keywords": [
"data pipeline",
" data preprocessing",
" data processing pipeline"
],
"urls": [
{
"comment_text": "",
"digests": {
"blake2b_256": "49cf2b8b0db58b912cd0314751c89bcfc93c536bbe0cd4de5426e0f7d7eceba7",
"md5": "772891811fb989f7a14e7904c36f4fee",
"sha256": "335b7347980a4268d12e0b5f90f21b76bd3bad2bc0cf3a4f69b3c29859d9b574"
},
"downloads": -1,
"filename": "discovery_capability-0.23.18-py3-none-any.whl",
"has_sig": false,
"md5_digest": "772891811fb989f7a14e7904c36f4fee",
"packagetype": "bdist_wheel",
"python_version": "py3",
"requires_python": ">=3.8",
"size": 6798403,
"upload_time": "2024-04-21T21:48:14",
"upload_time_iso_8601": "2024-04-21T21:48:14.505604Z",
"url": "https://files.pythonhosted.org/packages/49/cf/2b8b0db58b912cd0314751c89bcfc93c536bbe0cd4de5426e0f7d7eceba7/discovery_capability-0.23.18-py3-none-any.whl",
"yanked": false,
"yanked_reason": null
},
{
"comment_text": "",
"digests": {
"blake2b_256": "2ebb442f3976408645cc5c161c344827fc545ee95eb3d01fd17384ab05216458",
"md5": "7ae77414736214006bc09bef8530a615",
"sha256": "38b4da2d27dda7f8ee2b373e6382f14d8e199b51545795971ad77b48586249bb"
},
"downloads": -1,
"filename": "discovery-capability-0.23.18.tar.gz",
"has_sig": false,
"md5_digest": "7ae77414736214006bc09bef8530a615",
"packagetype": "sdist",
"python_version": "source",
"requires_python": ">=3.8",
"size": 6640190,
"upload_time": "2024-04-21T21:48:36",
"upload_time_iso_8601": "2024-04-21T21:48:36.232333Z",
"url": "https://files.pythonhosted.org/packages/2e/bb/442f3976408645cc5c161c344827fc545ee95eb3d01fd17384ab05216458/discovery-capability-0.23.18.tar.gz",
"yanked": false,
"yanked_reason": null
}
],
"upload_time": "2024-04-21 21:48:36",
"github": true,
"gitlab": false,
"bitbucket": false,
"codeberg": false,
"github_user": "gigas64",
"github_project": "discovery-capability",
"github_not_found": true,
"lcname": "discovery-capability"
}